
MiniMax Hailuo-02 Image To Video API (Standard, 768p, 512p): Advanced image-to-video generation model with 768p and 512p resolutions

Animates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.

Turns a single image into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count
![FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8689a8%2Fbbcmo6U5xg_RxDXijtxNA_55df705e1b1b4535a90bccd70887680e.jpg/tr:w-1920,q-80/bbcmo6U5xg_RxDXijtxNA_55df705e1b1b4535a90bccd70887680e.webp)
FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.

Professional video upscaling powered by Topaz Labs. Precision models (Proteus, Artemis, Iris, Dione, Theia, Gaia, Rhea) enhance footage up to 4x while staying faithful to the source. Best for clean, natural upscales of real-world footage.

Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.

Converts a given raster image to SVG format using Recraft model.

CassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. At 44.1 kHz stereo audio, expect a level of professional consistency with no breaks, no squeaks, and no random interruptions in your creations.

Gemini 3.1 Flash Image (a.k.a. Nano Banana 2) is Google's new state-of-the-art fast image generation and editing model

Kling 3.0 Turbo Standard animates a first and last frame reference image into 720P video with native audio, delivering quick, affordable image-driven motion for fast turnaround

Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.

Get encoding metadata from video and audio files using FFmpeg API.
![Image-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—in a flash.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a871484%2FfjLSktGKoWIGQWm-GRaUM_87cd94bbbff7400b830e73b8f6f075d4.jpg/tr:w-1920,q-80/fjLSktGKoWIGQWm-GRaUM_87cd94bbbff7400b830e73b8f6f075d4.webp)
Image-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—in a flash.

Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

Transform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.

Endpoint for Qwen's Image Editing 2511 model.

Frontier image editing model.

Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.

Image to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost
![Fastest inference in the world for the 12 billion parameter FLUX.1 [schnell] text-to-image model.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-2.jpg/tr:w-1920,q-80/Upscale-2.webp)
Fastest inference in the world for the 12 billion parameter FLUX.1 [schnell] text-to-image model.

SOTA stemming model for voice, drums, bass, guitar and more.

Kling 2.5 Turbo Standard: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

Professional generative image upscaling powered by Topaz Labs. Wonder 3.5 leads the range, with Redefine for prompt-guided detail and Recovery for extreme low-resolution sources. Best for rebuilding sharp detail in small or blurry images.

Image editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.

Generate high-quality 3D models from a single image using Tripo H3.1.

Train styles, people and other subjects at blazing speeds.

Kling 2.1 Pro is an advanced endpoint for the Kling 2.1 model, offering professional-grade videos with enhanced visual fidelity, precise camera movements, and dynamic motion control, perfect for cinematic storytelling.