
Generate videos from your image prompts using Veo 3.1 fast.

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
Generate sound effects using ElevenLabs advanced sound effects model.

H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space

Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

Clarity upscaler for upscaling images with high very fidelity.

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

Generate high quality, realistic music with fine controls using Elevenlabs Music!

bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long

Generate videos with audio with Seedance 1.5 (supports start & end frame)
![Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2Fpenguin%2FzSBCJtPpeIQwR5AC_IamX_b1e1137961754e4d851907c21f8c20cd.jpg/tr:w-1920,q-80/zSBCJtPpeIQwR5AC_IamX_b1e1137961754e4d851907c21f8c20cd.webp)
Text-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind

Generate videos from images with audio using xAI's Grok Imagine Video model.

Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model
Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.

Generate highly aesthetic images with xAI's Grok Imagine Image generation model.

Google's famous original image generation and editing model, a.k.a Nano Banana
![Image-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8a7f50%2FX8ffS5h55gcigsNZoNC7O_52e6b383ac214d2abe0a2e023f03de88.jpg/tr:w-1920,q-80/X8ffS5h55gcigsNZoNC7O_52e6b383ac214d2abe0a2e023f03de88.webp)
Image-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

Meta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.

Edit images with xAi's Grok Imagine 2.0 model.

Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.

Veo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video

Generate 3D models from images with Hunyuan 3D Pro

GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.

Kling 2.1 Standard is a cost-efficient endpoint for the Kling 2.1 model, delivering high-quality image-to-video generation