
Generate synced sounds for any video, and return the new sound track (like MMAudio)

Answer questions from the images.

A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.

Vidu Q1 Image to Video generates high-quality 1080p videos with exceptional visual quality and motion diversity from a single image

Juggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.

Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.

GOT-OCR2 works on a wide range of tasks, including plain document OCR, scene text OCR, formatted document OCR, and even OCR for tables, charts, mathematical formulas, geometric shapes, molecular formulas and sheet music.

Edit any image with a natural-language instruction using Bernini-R, changing the weather, materials, objects, or style while preserving the original composition.

Edit images with a text prompt using Emu 3.5 Image

HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

Generate long videos in 720p/30fps from images using LongCat Video

See how you or others might look at different ages, from younger to older, while preserving core facial features.

Generate 3D models from one or more images using ReconViaGen 0.5

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.

Enhance and refine portrait photos with improved clarity and detail.

Generate high-quality video with audio from text using LTX-2.3
![FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-5.jpeg/tr:w-1920,q-80/Upscale-5.webp)
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.

MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.

State-of-the-art open-source model in aesthetic quality

Use the latest Vidu Q2 models which much more better quality and control on your videos.

Clone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create speeches of yours!

A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.

Transforms images into comic book style