
Sana v1.5 4.8B is a powerful text-to-image model that generates ultra-high quality 4K images with remarkable detail.

Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.

A high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.

Generate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Finegrain Eraser removes objects—along with their shadows, reflections, and lighting artifacts—using only natural language, seamlessly filling the scene with contextually accurate content.

Edit images with a text prompt using Emu 3.5 Image

Generate video with audio from audio, text and images using LTX-2

Generate videos from prompts using LTX Video-0.9.5

Image-to-image editing with Step1X-Edit v2 from StepFun. Reasoning-enhanced modifications through a thinking–editing–reflection loop with MLLM world knowledge for abstract instruction comprehension.

MiDaS depth estimation preprocessor.
![FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-5.jpeg/tr:w-1920,q-80/Upscale-5.webp)
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

Generate high quality images from text prompts using CogView4. Longer text prompts will result in better quality images.

FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.

A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.

HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.

See how you or others might look at different ages, from younger to older, while preserving core facial features.

Generate long videos in 720p/30fps from images using LongCat Video
![FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a9f92d3%2FYI0vlnMufwkKs0eTTXmM7_UVyjAsaK.png/tr:w-1920,q-80/YI0vlnMufwkKs0eTTXmM7_UVyjAsaK.webp)
FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Texture an existing geometry mesh using a reference image with Hi3D.

Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.

Place your subject in any scene you imagine, from enchanted forests to urban settings, with professional composition and lighting

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.