
Kling Omni 3: Top-tier image-to-image with flawless consistency.

Generate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.

Generate production-quality lipsync from any audio using VEED's most advanced model yet.

Generate video clips from your images using Kling 1.6 (pro)

Gemini 3.1 Flash Image (a.k.a Nano Banana 2) is Google's new state-of-the-art fast image generation and editing model

Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.

Wan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency throughout each shot.

sync-3 image to video turns a single still into a talking character, and works with any illustration or animated frame paired with a voice track

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

Kokoro is a lightweight text-to-speech model that delivers comparable quality to larger models while being significantly faster and more cost-efficient.
![Experimental version of FLUX.1 Kontext [pro] with multi image handling capabilities](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-2.jpg/tr:w-1920,q-80/Upscale-2.webp)
Experimental version of FLUX.1 Kontext [pro] with multi image handling capabilities

Predict whether an image is NSFW or SFW.

A video understanding model to analyze video content and answer questions about what's happening in the video based on user prompts.

OpenAI's latest image generation and editing model: gpt-1-image.

FLUX LoRA Image-to-Image is a high-performance endpoint that transforms existing images using FLUX models, leveraging LoRA adaptations to enable rapid and precise image style transfer, modifications, and artistic variations.

Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization.

Wan-2.1 is a image-to-video model that generates high-quality videos with high visual quality and motion diversity from images

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

Wan-Animate Replace is a model that can integrate animated characters into reference videos, replacing the original character while preserving the scene’s lighting and color tone for seamless environmental integration.

Generate images from text and images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

Train a custom LoRA on your own images to teach Krea 2 a new subject, character, or style. Provide a set of training images (and an optional trigger word), and the trainer outputs LoRA weights you can use for inference with the Krea 2 LoRA endpoint.
![Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-3.jpeg/tr:w-1920,q-80/Upscale-3.webp)
Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.

Transform and edit existing images with text-guided instructions using the WAN 2.7 model for creative image manipulation.

OpenAI's latest image generation and editing model: gpt-1-image.

SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.

Endpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.

FLUX 3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.