
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video

Google's famous original image generation and editing model, a.k.a Nano Banana

Transfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.

Wan-2.2 Turbo image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.

Bria Eraser enables precise removal of unwanted objects from images while maintaining high-quality outputs. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us

ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs.

Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

Open source text-to-audio model.

Pixverse's latest V6 Model

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

Upscale videos with Bytedance's video upscaler.
![FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a868a0f%2FzL7LNUIqnPPhZNy_PtHJq_330f66115240460788092cb9523b6aba.jpg/tr:w-1920,q-80/zL7LNUIqnPPhZNy_PtHJq_330f66115240460788092cb9523b6aba.webp)
FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.

Merge videos with standalone audio files or audio from video files.

Kling AI Avatar v2 Pro: The premium endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.

LatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.

Kling LipSync is an audio-to-video model that generates realistic lip movements from audio input.

A audio understanding model to analyze audio content and answer questions about what's happening in the audio based on user prompts.

Wan 2.5 image-to-video model.

Generate videos with audio from text using Grok Imagine Video.

Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model

fal-ai/wan/v2.2-A14B/image-to-video

Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
Generate text from speech using ElevenLabs advanced speech-to-text model.

An endpoint for personalized image generation using Flux as per given description.