
Run Any Stable Diffusion model with customizable LoRA weights.

Generate text embeddings using OpenAI-compatible API. Access embedding models like text-embedding-3-small, text-embedding-3-large (OpenAI), and other embedding models available through OpenRouter. Drop-in replacement for the OpenAI embeddings API. Powered by OpenRouter.

Generate videos from prompts using LTX Video

SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.

Generate images with transparent backgrounds using Ideogram Transparent model

FLUX 3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.

Generate high quality video clips from text and image prompts using PixVerse v5.5

Extend videos with xAI's Grok Imagine video model

Vision reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts an image plus a prompt and returns text.

Generate video clips from your prompts using MiniMax model

Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.

Generate dubbed videos or audios using ElevenLabs Dubbing feature!

References into video with synchronized audio using MiniMax H3

Extract seamless tiling textures with PBR attribute maps from images

FFMPEG Untility for Extracting nth Frame

Start with a simple text input to create dynamic generations that defy expectations in up to 1080p. Experience better image clarity and crisper, sharper visuals.

Vidu's Q3 Turbo Model
![Image-to-image editing with FLUX.2 [klein] 4B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8a7f49%2FnKsGN6UMAi6IjaYdkmILC_e20d2097bb984ad589518cf915fe54b4.jpg/tr:w-1920,q-80/nKsGN6UMAi6IjaYdkmILC_e20d2097bb984ad589518cf915fe54b4.webp)
Image-to-image editing with FLUX.2 [klein] 4B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

Generate high-fidelity images extremely fast from text with Krea 2 Medium Turbo, supporting aspect ratio, creativity, seed controls, and optional style references.

Create seamless transition between images using PixVerse v5

Generate premium-quality images from text prompts using the enhanced WAN 2.7 Pro model with superior detail and composition.

Modify consistent characters while preserving their core identity. Edit poses, expressions, or clothing without losing recognizable character features

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

The OpenRouter Responses API with fal, powered by OpenRouter, provides unified access to a wide range of large language models - including GPT, Claude, Gemini, and many others through a single API interface.

Pixelcut's Video Background Remover is an AI segmentation model that erases backgrounds frame by frame, with seamless temporal consistency.
![FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2Ffor%2520videos-4.jpg/tr:w-1920,q-80/for%20videos-4.webp)
FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.

Place any product in any scenery with just a prompt or reference image while maintaining high integrity of the product. Trained exclusively on licensed data for safe and risk-free commercial use and optimized for eCommerce.

VOID removes objects from videos along with all interactions they induce on the scene