
HunyuanAvatar is a High-Fidelity Audio-Driven Human Animation model for Multiple Characters .

Kling 2.1 Master: The premium endpoint for Kling 2.1, designed for top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

Kling 2.1 Pro is an advanced endpoint for the Kling 2.1 model, offering professional-grade videos with enhanced visual fidelity, precise camera movements, and dynamic motion control, perfect for cinematic storytelling.

Kling 2.1 Standard is a cost-efficient endpoint for the Kling 2.1 model, delivering high-quality image-to-video generation

HunyuanPortrait is a diffusion-based framework for generating lifelike, temporally consistent portrait animations.

Generate video clips from your multiple image references using Kling 1.6 (standard)

Generate video clips from your multiple image references using Kling 1.6 (pro)

Generate videos from prompts and images using LTX Video-0.9.7 13B and custom LoRA

Generate videos from prompts and images using LTX Video-0.9.7 and custom LoRA

HunyuanCustom revolutionizes video generation with unmatched identity consistency across multiple input types. Its innovative fusion modules and alignment networks outperform competitors, maintaining subject integrity while responding flexibly to text, image, audio, and video conditions.

Deprecated. Use fal-ai/ltx-video-13b-dev or fal-ai/ltx-video-13b-distilled instead.

MAGI-1 generates videos from images with exceptional understanding of physical interactions and prompting

Generate video clips from your images using Kling 2.0 Master

Pika Effects are AI-powered video effects designed to modify objects, characters, and environments in a fun, engaging, and visually compelling manner.

Veo 2 creates videos from images with realistic motion and very high quality output.

Generate videos from prompts and images using LTX Video-0.9.5

SkyReels V1 is the first and most advanced open-source human-centric video foundation model. By fine-tuning HunyuanVideo on O(10M) high-quality film and television clips

Generate video clips from your images using Kling 1.6 (std)

Generate video clips from your images using Kling 1.6 (pro)

Generate video clips from your images using Kling 1.0

Generate video clips from your images using Kling 1.5 (pro)
fal is the best developer-friendly, one-stop shop for AI image-to-video models. Every image-to-video model on fal runs through the same SDK pattern, so once you’ve integrated one, switching between Seedance 2.0, Kling 3.0 Pro, or Veo 3.1 is a one-line endpoint change.
Image-to-video endpoints take an image URL plus a text prompt describing the motion, and return a URL to a generated video file. After installing @fal-ai/client and setting your FAL_KEY, it looks like this:
jsimport { fal } from "@fal-ai/client"; const result = await fal.subscribe("bytedance/seedance-2.0/image-to-video", { input: { prompt: "Slow cinematic push-in with the subject's hair moving gently in the wind", image_url: "https://your-host.com/photo.jpg", resolution: "720p" } }); console.log(result.data.video.url);
The same call shape works across Seedance 2.0, Kling 3.0 Pro, Veo 3.1, and the rest of the catalog. You swap the endpoint string and adjust the input fields each model expects. For example, Seedance 2.0 takes image_url, while Kling 3.0 Pro takes start_image_url.
Most premium image-to-video models on fal generate synchronized audio alongside the video itself.
generate_audio is on or off. Output covers music, lip-synced dialogue, and ambient sound.generate_audio is enabled, with support for multiple speakers and English and Chinese voice output. Audio adds 50% to the per-second cost, moving from $0.112 to $0.168.For projects where audio is part of the delivery, the cost difference can be meaningful. Enabling audio on Veo 3.1 doubles the rate, while Seedance bills the same either way.
Several models on fal accept an end frame alongside the start image, animating the transition between two specific points.
image_url and end_image_url. When both are provided, the model generates motion that transitions from the first frame to the second.veo3.1/first-last-frame-to-video and the fast variant, built for this workflow.end_image_url in the same way, with the same start-to-end transition behavior.For longer narratives, Kling 3.0 Pro’s multi_prompt feature lets you define multiple shots in sequence with distinct prompts and durations.
Pricing on fal generally scales by the second across image-to-video models, with resolution, audio, or both affecting the rate depending on the model.
| Model | Price |
|---|---|
| Kling 2.5 Turbo Pro | $0.35 for 5 seconds, then $0.07 / additional second |
| Kling 3.0 Pro | $0.112 / second audio off |
| Kling 3.0 Pro with audio | $0.168 / second |
| Kling 3.0 Pro with voice control | $0.196 / second |
| Seedance 2.0 Fast | $0.2419 / second at 720p, audio included |
| Seedance 2.0 Standard | $0.3024 / second at 720p, audio included |
| Veo 3.1 | $0.20 / second without audio at 720p or 1080p |
| Veo 3.1 with audio | $0.40 / second at 720p or 1080p |
As a worked example, a 5-second image-to-video clip costs roughly:
You only pay for what you generate, so you can compare motion quality, audio behavior, and frame-control options without rewriting your integration.
bashnpm install --save @fal-ai/client
bashexport FAL_KEY="YOUR_API_KEY"
jsimport { fal } from "@fal-ai/client"; const result = await fal.subscribe("bytedance/seedance-2.0/image-to-video", { input: { prompt: "Slow cinematic push-in with the subject's hair moving gently in the wind", image_url: "https://your-host.com/photo.jpg", resolution: "720p" } }); console.log(result.data.video.url);
The same auth, billing, and queue logic carry across every image-to-video endpoint, so you can compare models side by side without rewriting integration code.
For longer generations, higher-resolution outputs, or production workflows, submit to the queue and rely on webhooks instead of blocking on the result.