- Audio To Video
- Text To Video
Endpoint:
POST https://fal.run/fal-ai/kling-video/lipsync/audio-to-video
Endpoint ID: fal-ai/kling-video/lipsync/audio-to-videoTry it in the Playground
Run this model interactively with your own prompts.
Quick Start
Input Schema
string
required
The URL of the video to generate the lip sync for. Supports .mp4/.mov, ≤100MB, 2–10s, 720p/1080p only, width/height 720–1920px.
string
required
The URL of the audio to generate the lip sync for. Minimum duration is 2s and maximum duration is 60s. Maximum file size is 5MB.
Output Schema
File
required
The generated video
Input Example
Output Example
Performance
At $0.014 per 5-second video increment (rounded up), Kling LipSync positions as a specialized audio-to-video tool trading inference speed for lip-sync precision. Processing takes approximately 12 minutes regardless of video duration within the 2-10 second input range.Precision Lip-Sync Without Training Data
Kling LipSync uses audio-driven facial animation architecture that generates mouth movements directly from audio waveforms without requiring speaker-specific training data, contrasting with traditional lip-sync approaches that need extensive footage of the target speaker. What this means for you:- Zero-shot speaker adaptation: Sync any audio to any face without pre-training on that specific person, enabling rapid dubbing workflows across multiple speakers and languages
- Extended audio support: Process up to 60 seconds of audio against 2-10 second video clips, useful for looping background characters or extending dialogue beyond source footage length
- Format flexibility: Accepts 5 audio formats (.mp3, .wav, .ogg, .m4a, .aac) and standard video containers (.mp4, .mov), integrating into existing video generation workflows without format conversion overhead
- Increment-based pricing transparency: 5-second billing increments mean predictable costs (a 3-second video costs the same as 5 seconds at 0.028)
Technical Specifications
API Documentation | Quickstart Guide | Enterprise Pricing
Related
- Kling LipSync Text-to-Video — Video Generation
- Kling LipSync Audio-to-Video — Video Generation
Limitations
voice_languagerestricted to:zh,envoice_speedrange: 0.8 to 2