
Transfer expression from a video to a portrait.

Luma Uni-1 turns a text prompt into a single high-fidelity image, with control over aspect ratio and visual style, plus optional web-sourced and reference-image guidance for sharper grounding.

Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.

Restyle a video’s scene, lighting, and visual style from edited keyframes while preserving the source subjects’ identity, expressions, gaze, and motion. Developed by Eyeline Labs and Netflix researchers.

Wan 2.6 text-to-image model.

State of the art Multiview to 3D Object generation. Generate 3D models from multiple images!

Create high-quality images with accurate text rendering and rich knowledge details—supports editing, style transfer, and maintaining consistent characters across multiple images.

Generate images from text and a reference image using MiniMax Image-01 for consistent character appearance.

Moondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.

Modify a face to look younger or older while keeping identity realistic.

Luma Uni-1 Max generates a single image at the model's highest fidelity, delivering richer detail and stronger prompt adherence than the base tier for hero-quality stills.

FireRed Image Edit v1.1 is an updated version of FireRed Image Edit, with improved image editing capabilities.

Makes images more photorealistic and natural

Generate synced sounds for any video, and return the new sound track (like MMAudio)

Bria Extract Object uses text prompts to isolate a selected object from an image and return it as an RGBA PNG with a transparent background. Ideal for product, ecommerce, advertising, and creative editing workflows. Bria's Extract Object API leads in product shot extraction, outperforming SAM 3.1 where it counts most for commercial use.

Automatically retouches faces to smooth skin and remove blemishes.

ImagineArt 1.5 Pro is an advanced text-to-image model that creates ultra-high-fidelity 4K visuals with lifelike realism, refined aesthetics, and powerful creative output suited for professional use.

Predict poses from images.

Flux Vision Upscaler for magnify/upscaling images with high fidelity and creativity.

Generate videos from prompts using LTX Video-0.9.7 13B Distilled and custom LoRA
![Text-to-image generation with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8b09ab%2F3my7lbot7weIdE03-d5xc_2da235d3c4d14924b2c7a03f47e1bd65.jpg/tr:w-1920,q-80/3my7lbot7weIdE03-d5xc_2da235d3c4d14924b2c7a03f47e1bd65.webp)
Text-to-image generation with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.

Wan 2.5 image-to-image model.

Stable Diffusion 3 Medium (Text to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography, prompt understanding, and efficiency.

Change hairstyles and hair colors in photos realistically.