
Creates a reusable style from your reference images for use with Recraft V4 Styles Pro generation

Stable Cascade: Image generation on a smaller & cheaper latent space.

Generate video with audio from audio, text and images using LTX-2 Distilled

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

Wan-2.1 Pro is a premium text-to-video model that generates high-quality 1080p videos at 30fps with up to 6 seconds duration, delivering exceptional visual quality and motion diversity from text prompts

Place your subject in any scene you imagine, from enchanted forests to urban settings, with professional composition and lighting

Use USO to perform subject driven generations using reference image.

Run inference on LoRA adapters for TRELLIS.2 model
![Juggernaut Base Flux by RunDiffusion is a drop-in replacement for Flux [Dev] that delivers sharper details, richer colors, and enhanced realism, while instantly boosting LoRAs and LyCORIS with full compatibility.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Fjuggernaut-flux-base.webp/tr:w-1920,q-80/juggernaut-flux-base.webp)
Juggernaut Base Flux by RunDiffusion is a drop-in replacement for Flux [Dev] that delivers sharper details, richer colors, and enhanced realism, while instantly boosting LoRAs and LyCORIS with full compatibility.

MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.

Fast, low-latency text-to-image model with high-quality output and full JSON-structured controllability. Open-source, trained on licensed data, and optimized for production-scale generation.

Train a MiniMax H3 LoRA with first-frame conditioning, so a still image animates into video with audio; captions optional.

Structured Prompt Generation endpoint for Fibo, Bria's SOTA Open source model.

Generate video with audio from text using LTX-2.3 Distilled
Any pose, any style, any identity

M-LSD line segment detection preprocessor.

LoRA inference endpoint for the Qwen Image Editing model.

Generate videos from images and prompts using CogVideoX-5B

Train a MiniMax H3 LoRA on your own captioned clips for pure text-to-video generation with matching audio.
![FLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FTraining-5.jpg/tr:w-1920,q-80/Training-5.webp)
FLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

Texture an existing geometry mesh using a reference image with Hi3D.

Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their original endpoint via causal continuation.

Accelerated image generation with Ideogram V2 Turbo. Create high-quality visuals, posters, and logos with enhanced speed while maintaining Ideogram's signature quality.

Train custom LoRAs for personalization, styles or other use cases on top of Ideogram V4.