
Edit images with a text prompt using Emu 3.5 Image

A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.

Vidu Reference to Video creates videos by using a reference images and combining them with a prompt.

Finegrain Eraser removes any object selected with a bounding box—along with its shadows, reflections, and lighting artifacts—seamlessly reconstructing the scene with contextually accurate content.

Generate long videos from text using LongCat Video Distilled
![FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-5.jpeg/tr:w-1920,q-80/Upscale-5.webp)
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

Kandinsky 5.0 Distilled is a lightweight diffusion model for fast, high-quality text-to-video generation.

Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.

Recraft V3 Create Style is capable of creating unique styles for Recraft V3 based on your images.

A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency.

Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.

Generate video clips from your prompts using Kling 1.6 (std)

Lumina-Image-2.0 is a 2 billion parameter flow-based diffusion transforer which features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.

Image generation with BitDance. Fast, high-resolution photorealistic images using an autoregressive LLM— for efficient, high-quality results.

Transforms images into comic book style

Generate HDR from reference video using LTX-2.3

Interpolate between video frames

MiDaS depth estimation preprocessor.

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.

Vidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image

Generate 3D models from multiple view images using Hi3D.

Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

Generate synced sounds for any video, and return the new sound track (like MMAudio)