LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.
LTX logo
lightricks/ltx-2.5/image-to-video/fast

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.

stylized
transform
lip-sync
image-to-video
Generate speech with expressive and realistic voices from xAI
xAI logo
xai/tts/v1

Generate speech with expressive and realistic voices from xAI

text-to-speech
Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
sonilo/v1.1/text-to-sound-effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

sfx
audio
effects
text-to-audio
Wan 2.6 image-to-video model.
Alibaba logo
wan/v2.6/image-to-video

Wan 2.6 image-to-video model.

image-to-video
Kling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
Kling logo
kling-video/o3/pro/video-to-video/reference

Kling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

video-to-video
Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
Kling logo
kling-video/o1/image-to-video

Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.

image-to-video
Merge audios into a single audio using FFmpeg API!
ffmpeg-api/merge-audios

Merge audios into a single audio using FFmpeg API!

ffmpeg
audio-to-audio
Edits generated video across multiple conversational turns while preserving scene coherence. Applies iterative changes through natural-language instructions without regenerating the full sequence from scratch.
Google logo
google/gemini-omni-flash/edit

Edits generated video across multiple conversational turns while preserving scene coherence. Applies iterative changes through natural-language instructions without regenerating the full sequence from scratch.

stylized
transform
lipsync
video-to-video
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
inworld-tts

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

inworld
tts
text-to-speech
Professional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.
Topaz Labs logo
topaz/upscale/video/generative

Professional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.

upscale
video
video-to-video
Generate video clips from your prompts using Kling 1.6 (std)
Kling logo
kling-video/v1.6/standard/text-to-video

Generate video clips from your prompts using Kling 1.6 (std)

text-to-video
Pixverse's latest v6 Model.
Pixverse logo
pixverse/v6/text-to-video

Pixverse's latest v6 Model.

text-to-video
Alibaba's #1-ranked Happy Horse 1.0 — generate 1080p video with synchronized native audio and multilingual lip-sync from text prompts or images.
Alibaba logo
alibaba/happy-horse/image-to-video

Alibaba's #1-ranked Happy Horse 1.0 — generate 1080p video with synchronized native audio and multilingual lip-sync from text prompts or images.

video
happy-horse
image-to-video
Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.
ElevenLabs logo
elevenlabs/text-to-dialogue/eleven-v3

Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.

audio
text-to-audio
Generate realistic videos using Kling O3 from Kling Team!
Kling logo
kling-video/o3/pro/text-to-video

Generate realistic videos using Kling O3 from Kling Team!

text-to-video
MiniMax Hailuo-2.3-Fast Image To Video API (Standard, 768p): Advanced fast image-to-video generation model with 768p resolution
Minimax logo
minimax/hailuo-2.3-fast/standard/image-to-video

MiniMax Hailuo-2.3-Fast Image To Video API (Standard, 768p): Advanced fast image-to-video generation model with 768p resolution

image-to-video
Qwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers.
Alibaba logo
qwen-image-layered

Qwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers.

qwen
layer
image-to-image
Text-to-Image endpoint with LoRA support for Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
Alibaba logo
z-image/turbo/lora

Text-to-Image endpoint with LoRA support for Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

z-image
lora
fast
text-to-image
Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.
sam-audio/separate

Audio separation with SAM Audio. Isolate any sound using natural language—professional-grade audio editing made simple for creators, researchers, and accessibility applications.

sam-audio
audio-to-audio
Edit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
Kling logo
kling-video/o1/video-to-video/edit

Edit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.

video-to-video
Run any video-capable LLM with fal. Analyze, summarize, and understand video files using Gemini (Google) models. Supports mp4, mpeg, mov, webm, and YouTube links. Powered by OpenRouter.
openrouter/router/video

Run any video-capable LLM with fal. Analyze, summarize, and understand video files using Gemini (Google) models. Supports mp4, mpeg, mov, webm, and YouTube links. Powered by OpenRouter.

video-to-text
Generate 3D models from multiple view images using Tripo H3.1.
tripo3d/h3.1/multiview-to-3d

Generate 3D models from multiple view images using Tripo H3.1.

3d
multiview-to-3d
3d-generation
image-to-3d
Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.
Bria AI logo
bria/video/background-removal/v3

Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.

video-to-video
A versatile endpoint for the FLUX.1 [dev] model that supports multiple AI extensions including LoRA, ControlNet conditioning, and IP-Adapter integration, enabling comprehensive control over image generation through various guidance methods.
Black Forest Labs logo
flux-general

A versatile endpoint for the FLUX.1 [dev] model that supports multiple AI extensions including LoRA, ControlNet conditioning, and IP-Adapter integration, enabling comprehensive control over image generation through various guidance methods.

lora
controlnet
ip-adapter
text-to-image
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
Bytedance logo
bytedance/seedance-2.0/mini/text-to-video

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

stylized
transform
lipsync
text-to-video
Generates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.
depth-anything-video

Generates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.

video to video
motion
edit
video-to-video
MiniMax Hailuo-02 Image To Video API (Pro, 1080p): Advanced image-to-video generation model with 1080p resolution
Minimax logo
minimax/hailuo-02/pro/image-to-video

MiniMax Hailuo-02 Image To Video API (Pro, 1080p): Advanced image-to-video generation model with 1080p resolution

image-to-video
Edit and transform images using text instructions with the WAN 2.7 Pro model for precise, professional-grade image modifications.
Alibaba logo
wan/v2.7/pro/edit

Edit and transform images using text instructions with the WAN 2.7 Pro model for precise, professional-grade image modifications.

wan
image-editing
pro
image-to-image
Showing 281 to 308 of 1504 results