Try H3 Max Director
Log-inSign-up
Contact SalesLog-inSign-up
  • Documentation
  • Pricing
  • Enterprise
Loading Explore
Search models...Search by model, task, category and more
View all models
Search models...Search by model, task, category and more View all models
Try:
Newest image to video models
Flux Kontext
Generate 3D model
Create music
Remove background
Upscale
Training
Try on clothing

Ready to transform your enterprise with AI?

Contact Sales

Learn more

StatusAbout UsDocumentationTrust & SafetyVerify fal-Generated ContentCareersPricingBlogEnterpriseGet in touchReport ContentGrantsEventsLegalSite DirectoryPressLearnGen Media Report Vol. 2

Image Models

Seedream 5.0GPT Image 2.5GPT Image 2Flux 2Nano Banana 2.1Nano Banana 2Ideogram 4Krea 2Nano Banana ProQwen Image 3Explore More

Video Models

AI Video GeneratorText to VideoSeedance 2.5Seedance 2.0Gemini OmniMiniMax H3Kling 3.0Veo 3.1Grok Imagine 1.5Happy Horse 1.1Happy OysterWan 3.0LTX 2.3PixVerse V6

Labs

Black Forest LabsGoogleOpenAIxAIAlibabaKlingByteDanceElevenLabs

Playgrounds

fal AgentSandboxWorkflowsTrainingFree ToolsBackground RemoverImage UpscalerImage ExtenderImage Resizer

Socials

DiscordGitHubRedditTwitterLinkedInYouTubeInstagramTikTok
Features and Labels, 2026. All Rights Reserved. Terms of Service and Privacy Policy.
MiniMax H3 Max Text to Video
text-to-video
stylized
transform
lipsync

MiniMax H3 Max Text to Video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Try it now!See docs
FLUX 3
image-to-video
stylized
transform
lipsync

FLUX 3

FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

Try it now!See docs
image-to-video
stylized
transform
lipsync

Seedance 2.5 Image to Video

Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.

Try it now!See docs
image-to-video
stylized
transform
lipsync

MiniMax H3

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

Try it now!See docs
image-to-video
stylized
transform
lipsync

MiniMax H3 Max Image to Video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Try it now!See docs
image-to-video
image-to-video

Kling Video v3 Image to Video [Pro]

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

Try it now!See docs
text-to-image
stylized
transform
typography

Krea 2 Turbo

Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.

Try it now!See docs
Search models...Search by model, task, category and more
View all models
Search models...Search by model, task, category and more View all models
Try:
Newest image to video models
Flux Kontext
Generate 3D model
Create music
Remove background
Upscale
Training
Try on clothing

Trending

Models that are popular with developers right now.

Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro/edit

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
image-to-image
Nano Banana 2 is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-2/edit

Nano Banana 2 is Google's new state-of-the-art image generation and editing model

image-to-image
Editing built for the tightest control, edits scoped precisely to the instruction, with subject and composition preserved across many rounds of revision.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/sunburst/edit

Editing built for the tightest control, edits scoped precisely to the instruction, with subject and composition preserved across many rounds of revision.

stylized
transform
editing
image-to-image
Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/schnell

FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.

text-to-image
GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.
OpenAI logoOpenAI logo
openai/gpt-image-2/edit

GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.

gpt-image-2
openai
chatgpt-images-2
image-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
Precise image editing that changes only what's asked, keeping subject, composition, and background intact, with reference subjects staying recognizable across styles and successive edits.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/flare/edit

Precise image editing that changes only what's asked, keeping subject, composition, and background intact, with reference subjects staying recognizable across styles and successive edits.

stylized
transform
editing
image-to-image

Recently Added

Newly added models across image, video, audio, and more.

Relight any video with H3 Max from a lighting sphere image, while preserving the source subjects, motion, camera and audio.
new
Minimax logo
minimax/h3-max/relight

Relight any video with H3 Max from a lighting sphere image, while preserving the source subjects, motion, camera and audio.

video-editing
relighting
video-to-video
Vidu Q4 turns a first-frame image into 3 to 16 second videos with native audio at up to 4K. Best for animating stills, product shots and character scenes with dialogue.
new
vidu/q4/image-to-video

Vidu Q4 turns a first-frame image into 3 to 16 second videos with native audio at up to 4K. Best for animating stills, product shots and character scenes with dialogue.

multi-shot
4k
image-to-video
Vidu Q4 generates videos from up to 12 reference images and 3 voice clips, keeping characters, objects, and voices consistent across up to 16 seconds at up to 4K.
new
vidu/q4/reference-to-video

Vidu Q4 generates videos from up to 12 reference images and 3 voice clips, keeping characters, objects, and voices consistent across up to 16 seconds at up to 4K.

reference-to-video
multi-shot
4k
image-to-video
Generates podcast audio from an ordered sequence of conversation turns, with text and a voice selection for each turn.
new
mureka/api/generate/podcast

Generates podcast audio from an ordered sequence of conversation turns, with text and a voice selection for each turn.

podcast
speech
audio generation
text-to-audio
Generates songs from a text prompt or supplied lyrics, with musical style, vocal, reference-audio, and melody controls, plus support for multiple variations.
new
mureka/api/generate/song

Generates songs from a text prompt or supplied lyrics, with musical style, vocal, reference-audio, and melody controls, plus support for multiple variations.

music
vocals
song generation
text-to-audio
Generates instrumental music from a text description or uploaded instrumental reference, with support for multiple variations.
new
mureka/api/generate/instrumental

Generates instrumental music from a text description or uploaded instrumental reference, with support for multiple variations.

music
instrumental
music generation
text-to-audio
Writes song lyrics and a title from a text prompt describing a topic, theme, or song concept.
new
mureka/api/generate/lyrics

Writes song lyrics and a title from a text prompt describing a topic, theme, or song concept.

music
lyrics
text generation
text-to-text
Creates a lyrics video from generated or uploaded audio, with selectable layouts, aspect ratios, backgrounds, and lyric-row or time-range selection.
new
mureka/api/generate/lyrics-video

Creates a lyrics video from generated or uploaded audio, with selectable layouts, aspect ratios, backgrounds, and lyric-row or time-range selection.

music
lyrics
video
audio-to-video

Model Labs

Explore the AI labs powering models on fal

Decart logo
Decart
Kling logoKling logo
Kling
Minimax logo
Minimax
Topaz Labs logo
Topaz Labs
LTX logo
LTX
xAI logoxAI logo
xAI
OpenAI logoOpenAI logo
OpenAI
Krea logoKrea logo
Krea
ElevenLabs logoElevenLabs logo
ElevenLabs
Bytedance logoBytedance logo
Bytedance
Alibaba logoAlibaba logo
Alibaba
Google logoGoogle logo
Google
Bria AI logoBria AI logo
Bria AI
Black Forest Labs logoBlack Forest Labs logo
Black Forest Labs
Veed logoVeed logo
Veed
Heygen logo
Heygen
Luma AI logoLuma AI logo
Luma AI
Ideogram logoIdeogram logo
Ideogram
Pixverse logo
Pixverse

H3 Max by fal

H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.
new
Minimax logo
minimax/h3-max/lip-sync/image-to-video

H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.

lipsync
animation
audio
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/reference-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
typography
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/image-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/text-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space
new
Minimax logo
minimax/h3-max/camera-controls

H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space

stylized
transform
editing
image-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/text-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
Direct continuous, realtime video streams with live prompts while preserving characters, settings, and story continuity.
Minimax logo
minimax/h3-max/director

Direct continuous, realtime video streams with live prompts while preserving characters, settings, and story continuity.

text-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/image-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video

New and Noteworthy

State-of-the-art models we think you'll love!

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.
Alibaba logoAlibaba logo
alibaba/wan-3.0-prime/image-to-video

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.

image
video
image-to-video
MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.
Minimax logo
minimax/h3/image-to-video

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

stylized
transform
lipsync
image-to-video
LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.
LTX logo
lightricks/ltx-2.5/image-to-video/fast

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.

stylized
transform
lip-sync
image-to-video
Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence
Alibaba logoAlibaba logo
alibaba/qwen-image-3/text-to-image

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence

stylized
transform
typography
text-to-image
FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.
Black Forest Labs logoBlack Forest Labs logo
blackforestlabs/flux-3/image-to-video

FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

stylized
transform
lipsync
image-to-video
Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.
xAI logoxAI logo
xai/grok-imagine-video/v1.5/image-to-video

Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.

stylized
transform
lipsync
image-to-video
Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
Kling logoKling logo
kling-video/v3/pro/image-to-video

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video
Nano Banana 2 is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-2/edit

Nano Banana 2 is Google's new state-of-the-art image generation and editing model

image-to-image

Seedance 2.5

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.
Bytedance logoBytedance logo
bytedance/seedance-2.5/text-to-video

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.

stylized
transform
lipsync
text-to-video
Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.
Bytedance logoBytedance logo
bytedance/seedance-2.5/image-to-video

Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.

stylized
transform
lipsync
image-to-video
Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.
Bytedance logoBytedance logo
bytedance/seedance-2.5/reference-to-video

Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.

stylized
transform
lipsync
image-to-video

Video models

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/reference-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
typography
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/image-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/text-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/image-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/text-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
Kling logoKling logo
kling-video/v2.5-turbo/pro/image-to-video

Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

stylized
transform
image-to-video
Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
Kling logoKling logo
kling-video/v3/pro/image-to-video

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video
Kling 3.0 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.
Kling logoKling logo
kling-video/v3/pro/text-to-video

Kling 3.0 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.

text-to-video

Image models

Nano Banana 2 is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-2/edit

Nano Banana 2 is Google's new state-of-the-art image generation and editing model

image-to-image
FLUX 3 Image Edit from Black Forest Labs makes precise local edits without changing the rest of the image, and combines up to 10 references into one balanced, well-composed result.
new
Black Forest Labs logoBlack Forest Labs logo
blackforestlabs/flux-3/edit-image

FLUX 3 Image Edit from Black Forest Labs makes precise local edits without changing the rest of the image, and combines up to 10 references into one balanced, well-composed result.

flux-3-image
black-forest-labs
multi-reference
image-to-image
GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.
OpenAI logoOpenAI logo
openai/gpt-image-2/edit

GPT Image 2, OpenAI's latest image model, is capable of making fine-grained, detailed edits to images.

gpt-image-2
openai
chatgpt-images-2
image-to-image
GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.
OpenAI logoOpenAI logo
openai/gpt-image-2

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.

gpt-image-2
openai
typography
text-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro/edit

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
image-to-image
Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
Google's famous original image generation and editing model
Google logoGoogle logo
nano-banana/edit

Google's famous original image generation and editing model

image-editing
image-to-image

Audio models

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.
ElevenLabs logoElevenLabs logo
elevenlabs/tts/multilingual-v2

Generate multilingual text-to-speech audio using ElevenLabs TTS Multilingual v2.

audio
text-to-audio
Generate text-to-speech audio using Eleven-v3 from ElevenLabs.
ElevenLabs logoElevenLabs logo
elevenlabs/tts/eleven-v3

Generate text-to-speech audio using Eleven-v3 from ElevenLabs.

audio
text-to-audio
Generate sound effects using ElevenLabs advanced sound effects model.
ElevenLabs logoElevenLabs logo
elevenlabs/sound-effects/v2

Generate sound effects using ElevenLabs advanced sound effects model.

sound
text-to-audio
Generate music from a simple prompt using ACE-Step
ace-step/prompt-to-audio

Generate music from a simple prompt using ACE-Step

text-to-music
text-to-audio
Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.
ElevenLabs logoElevenLabs logo
elevenlabs/tts/turbo-v2.5

Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.

audio
text-to-speech
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Minimax logo
minimax/speech-02-hd

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
text-to-speech
Generate speech with expressive and realistic voices from xAI
xAI logoxAI logo
xai/tts/v1

Generate speech with expressive and realistic voices from xAI

text-to-speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Minimax logo
minimax/speech-2.8-hd

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

text-to-speech

3D models

SAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.
sam-3/3d-objects

SAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.

3d
object
image-to-3d
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
hunyuan-motion

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

motion
text-to-3d
Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.
trellis

Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-3d
Generate high-quality 3D models from a single image using Tripo H3.1.
tripo3d/h3.1/image-to-3d

Generate high-quality 3D models from a single image using Tripo H3.1.

3d
3d-generation
tripo
image-to-3d
Generate 3D models from text descriptions using Tripo H3.1.
tripo3d/h3.1/text-to-3d

Generate 3D models from text descriptions using Tripo H3.1.

3d
3d-generation
tripo
text-to-3d
Generate 3D models from multiple view images using Tripo H3.1.
tripo3d/h3.1/multiview-to-3d

Generate 3D models from multiple view images using Tripo H3.1.

3d
multiview-to-3d
3d-generation
image-to-3d
Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.
meshy/v6-preview/image-to-3d

Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.

image-to-3d
Generate 3D models from images with Hunyuan 3D Pro
hunyuan-3d/v3.1/pro/image-to-3d

Generate 3D models from images with Hunyuan 3D Pro

3d
hunyuan
image-to-3d

Grok Imagine

Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.
xAI logoxAI logo
xai/grok-imagine-video/v1.5/image-to-video

Generate videos from images with audio using xAI's Grok Imagine 1.5 Video model.

stylized
transform
lipsync
image-to-video
Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.
xAI logoxAI logo
xai/grok-imagine-video/v1.5/text-to-video

Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.

stylized
transform
lipsync
text-to-video
Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.
xAI logoxAI logo
xai/grok-imagine-video/v1.5/reference-to-video

Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.

stylized
transform
lipsync
image-to-video
Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.
xAI logoxAI logo
xai/grok-imagine-image/quality/edit

Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.

stylized
transform
typography
image-to-image
Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.
xAI logoxAI logo
xai/grok-imagine-image/quality/text-to-image

Grok Imagine Pro is an advanced AI model from xAI that creates high-quality visuals from text prompts and allows you to edit or analyze existing images.

stylized
transform
typography
text-to-image
Extend videos with xAI's Grok Imagine video model
xAI logoxAI logo
xai/grok-imagine-video/extend-video

Extend videos with xAI's Grok Imagine video model

video-edit
v2v
grok
video-to-video
Generate speech with expressive and realistic voices from xAI
xAI logoxAI logo
xai/tts/v1

Generate speech with expressive and realistic voices from xAI

text-to-speech
Edit videos using xAI's Grok Imagine
xAI logoxAI logo
xai/grok-imagine-video/edit-video

Edit videos using xAI's Grok Imagine

video-edit
v2v
grok
video-to-video

Best AI Image Generators

Unlock the future of creativity with these text to image, AI image generator models.

Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/schnell

FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.

text-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.
OpenAI logoOpenAI logo
openai/gpt-image-2

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.

gpt-image-2
openai
typography
text-to-image
OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/flare/text-to-image

OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.

realism
typography
stylized
text-to-image
FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/dev

FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.

text-to-image
OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/sunburst/text-to-image

OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.

realism
typography
stylized
text-to-image
Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows
Black Forest Labs logoBlack Forest Labs logo
flux-2-pro

Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows

text-to-image

Best Image Editing Models

The fan favorite best image editing models on the market

Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro/edit

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
image-to-image
A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.
Bytedance logoBytedance logo
bytedance/seedream/v4/edit

A new-generation image creation model ByteDance, Seedream 4.0 integrates image generation and image editing capabilities into a single, unified architecture.

stylized
transform
editing
image-to-image
Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.
Bria AI logoBria AI logo
bria/fibo-edit-1.5/edit

Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.

stylized
transform
editing
image-to-image
FLUX.1 Kontext [pro] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.
Black Forest Labs logoBlack Forest Labs logo
flux-pro/kontext

FLUX.1 Kontext [pro] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.

image-to-image
Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.
Black Forest Labs logoBlack Forest Labs logo
flux-kontext-lora

Fast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.

image-editing
image-to-image

Best of Open Source

Some of our favorite open source media models

MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
Minimax logo
minimax/h3/reference-to-video

MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.

stylized
transform
lipsync
image-to-video
MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.
Minimax logo
minimax/h3/image-to-video

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

stylized
transform
lipsync
image-to-video
MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.
Minimax logo
minimax/h3/text-to-video

MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.

stylized
transform
lipsync
text-to-video
LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.
LTX logo
lightricks/ltx-2.5/image-to-video/fast

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint animates a still image into video with synchronized audio in a single pass, in a speed-optimized mode for quick iteration.

stylized
transform
lip-sync
image-to-video
Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.
Bria AI logoBria AI logo
bria/fibo-edit-1.5/edit

Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.

stylized
transform
editing
image-to-image
LoRA trainer for FLUX.1 Kontext [dev]
Black Forest Labs logoBlack Forest Labs logo
flux-kontext-trainer

LoRA trainer for FLUX.1 Kontext [dev]

training
Generate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA
LTX logo
ltx-video-13b-distilled/image-to-video

Generate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA

video
ltx-video
image-to-video
Wan 2.2 text to image LoRA trainer. Fine-tune Wan 2.2 for subjects and styles with unprecedented detail.
Alibaba logoAlibaba logo
wan-22-image-trainer

Wan 2.2 text to image LoRA trainer. Fine-tune Wan 2.2 for subjects and styles with unprecedented detail.

lora
personalization
training

Seedance 2.0

The new sota video model by Bytedance. Access the new stunning video generation model today.

ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.
Bytedance logoBytedance logo
bytedance/seedance-2.0/text-to-video

ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.

stylized
transform
lipsync
text-to-video
ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.
Bytedance logoBytedance logo
bytedance/seedance-2.0/reference-to-video

ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.

stylized
transform
lipsync
image-to-video
ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.
Bytedance logoBytedance logo
bytedance/seedance-2.0/fast/text-to-video

ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.

stylized
transform
lipsync
text-to-video
ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.
Bytedance logoBytedance logo
bytedance/seedance-2.0/image-to-video

ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.

stylized
transform
lipsync
image-to-video
ByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.
Bytedance logoBytedance logo
bytedance/seedance-2.0/fast/reference-to-video

ByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.

stylized
transform
lipsync
image-to-video
ByteDance's most advanced image-to-video model, fast tier. Lower latency and cost with synchronized audio, start and end frame control, and motion prompts.
Bytedance logoBytedance logo
bytedance/seedance-2.0/fast/image-to-video

ByteDance's most advanced image-to-video model, fast tier. Lower latency and cost with synchronized audio, start and end frame control, and motion prompts.

stylized
transform
lipsync
image-to-video

Text To Speech APIs

Create lifelike speech with our AI text to speech APIs

Generate text-to-speech audio using Eleven-v3 from ElevenLabs.
ElevenLabs logoElevenLabs logo
elevenlabs/tts/eleven-v3

Generate text-to-speech audio using Eleven-v3 from ElevenLabs.

audio
text-to-audio
Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.
Bytedance logoBytedance logo
bytedance/seed-audio-1.0

Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.

text-to-audio
Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
Google logoGoogle logo
gemini-3.1-flash-tts

Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

lipsync
avatar
text-to-speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Minimax logo
minimax/speech-2.8-hd

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

text-to-speech
Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library
async/tts-pro/v1.0

Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library

voice-clone
lipsync
text-to-speech
Generate speech with expressive and realistic voices from xAI
xAI logoxAI logo
xai/tts/v1

Generate speech with expressive and realistic voices from xAI

text-to-speech
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
Alibaba logoAlibaba logo
qwen-3-tts/text-to-speech/1.7b

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

text-to-speech
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
inworld-tts

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

inworld
tts
text-to-speech

AI Image Generator APIs

Generate a variety of stunning images using our AI Image Generator APIs

Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/schnell

FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.

text-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.
OpenAI logoOpenAI logo
openai/gpt-image-2

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.

gpt-image-2
openai
typography
text-to-image
OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/flare/text-to-image

OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.

realism
typography
stylized
text-to-image
FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/dev

FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.

text-to-image
OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/sunburst/text-to-image

OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.

realism
typography
stylized
text-to-image
Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows
Black Forest Labs logoBlack Forest Labs logo
flux-2-pro

Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows

text-to-image

Text to Video APIs

Access the top Text to Video APIs with lightning fast inference speeds

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/text-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.
Bytedance logoBytedance logo
bytedance/seedance-2.5/text-to-video

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.

stylized
transform
lipsync
text-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/text-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
text-to-video
Kling 3.0 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.
Kling logoKling logo
kling-video/v3/pro/text-to-video

Kling 3.0 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.

text-to-video
Faster and more cost effective version of Google's Veo 3.1!
Google logoGoogle logo
veo3.1/fast

Faster and more cost effective version of Google's Veo 3.1!

text-to-video
Kling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.
Kling logoKling logo
kling-video/v3/standard/text-to-video

Kling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.

text-to-video
ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.
Bytedance logoBytedance logo
bytedance/seedance-2.0/text-to-video

ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.

stylized
transform
lipsync
text-to-video
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
Alibaba logoAlibaba logo
alibaba/wan-3.0/text-to-video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylized
transform
lipsync
text-to-video

Image to Video APIs

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
Kling logoKling logo
kling-video/v3/pro/image-to-video

Kling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/image-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max-turbo/image-to-video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
lipsync
image-to-video
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Minimax logo
minimax/h3-max/reference-to-video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

stylized
transform
typography
image-to-video
Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.
Bytedance logoBytedance logo
bytedance/seedance-2.5/reference-to-video

Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.

stylized
transform
lipsync
image-to-video
Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
Kling logoKling logo
kling-video/v2.5-turbo/pro/image-to-video

Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

stylized
transform
image-to-video
Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.
Bytedance logoBytedance logo
bytedance/seedance-2.5/image-to-video

Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.

stylized
transform
lipsync
image-to-video
Kling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
Kling logoKling logo
kling-video/v3/standard/image-to-video

Kling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.

image-to-video

Best Image Models

Top-performing models for high-quality image generation and editing.

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.
OpenAI logoOpenAI logo
openai/gpt-image-2

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.

gpt-image-2
openai
typography
text-to-image
Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.
recraft/v4/pro/text-to-image

Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.

text-to-image
Text-to-image model with high-fidelity outputs, accurate typography, and style preset, strong in photorealism, textures, and beyond. JSON-structured prompts give enterprise and agentic workflows production-ready control. Trained on licensed data.
Bria AI logoBria AI logo
bria/fibo-gen-1.5/text-to-image

Text-to-image model with high-fidelity outputs, accurate typography, and style preset, strong in photorealism, textures, and beyond. JSON-structured prompts give enterprise and agentic workflows production-ready control. Trained on licensed data.

stylized
transform
realism
text-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
FLUX.1 Kontext [pro] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.
Black Forest Labs logoBlack Forest Labs logo
flux-pro/kontext

FLUX.1 Kontext [pro] handles both text and reference images as inputs, seamlessly enabling targeted, local edits and complex transformations of entire scenes.

image-to-image
Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.
Black Forest Labs logoBlack Forest Labs logo
flux-krea-lora/stream

Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

lora
personalization
text-to-image
Recraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.
recraft/v3/text-to-image

Recraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.

vector
typography
style
text-to-image

Background Remover APIs

Find the API of your choice to remove a background from your image or video

Pixelcut’s Background Remover enables fast, ultra high-quality removal of backgrounds from images. Perfect for e-commerce and image editing workflows. Powered by advanced AI for clean, perfect cutouts every time.
pixelcut/background-removal

Pixelcut’s Background Remover enables fast, ultra high-quality removal of backgrounds from images. Perfect for e-commerce and image editing workflows. Powered by advanced AI for clean, perfect cutouts every time.

background removal
utility
remove background
image-to-image
Remove backgrounds from existing images with Ideogram's remove background feature. Isolate subjects cleanly for compositing and creative reuse.
Ideogram logoIdeogram logo
ideogram/remove-background

Remove backgrounds from existing images with Ideogram's remove background feature. Isolate subjects cleanly for compositing and creative reuse.

image-to-image
bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
birefnet/v2

bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

background removal
segmentation
high-res
image-to-image
Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.
Bria AI logoBria AI logo
bria/video/background-removal/v3

Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.

video-to-video
Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
Bria AI logoBria AI logo
bria/video/background-removal/green-screen-despill

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

video-to-video
Remove the background from an image.
imageutils/rembg

Remove the background from an image.

background removal
utility
editing
image-to-image
Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04
Bria AI logoBria AI logo
bria/background/remove

Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04

background removal
image segmentation
high resolution
image-to-image
bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
birefnet

bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)

background removal
segmentation
high-res
image-to-image

Veo 3.1

Generate videos from a first/last frame using Google's Veo 3.1 Fast
Google logoGoogle logo
veo3.1/fast/first-last-frame-to-video

Generate videos from a first/last frame using Google's Veo 3.1 Fast

image-to-video
Generate videos from a first and last framed using Google's Veo 3.1
Google logoGoogle logo
veo3.1/first-last-frame-to-video

Generate videos from a first and last framed using Google's Veo 3.1

image-to-video
Faster and more cost effective version of Google's Veo 3.1!
Google logoGoogle logo
veo3.1/fast

Faster and more cost effective version of Google's Veo 3.1!

text-to-video
Generate videos from your image prompts using Veo 3.1 fast.
Google logoGoogle logo
veo3.1/fast/image-to-video

Generate videos from your image prompts using Veo 3.1 fast.

image-to-video
Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind
Google logoGoogle logo
veo3.1/image-to-video

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind

image-to-video
Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!
Google logoGoogle logo
veo3.1

Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!

text-to-video
Generate Videos from images using Google's Veo 3.1
Google logoGoogle logo
veo3.1/reference-to-video

Generate Videos from images using Google's Veo 3.1

image-to-video

Marquee Video Models

Flagship video generation models known for top-tier quality, motion control, and cinematic results.

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.
Alibaba logoAlibaba logo
alibaba/wan-3.0-prime/image-to-video

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.

image
video
image-to-video
Images into video with synchronized audio using MiniMax H3; your image becomes the first frame, prompt optional, with trained LoRA support for subject consistency.
Minimax logo
minimax/h3/image-to-video/lora

Images into video with synchronized audio using MiniMax H3; your image becomes the first frame, prompt optional, with trained LoRA support for subject consistency.

utility
editing
image-to-video
Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
Kling logoKling logo
kling-video/o3/standard/image-to-video

Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.

image-to-video
Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
Kling logoKling logo
kling-video/v2.5-turbo/pro/image-to-video

Kling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

stylized
transform
image-to-video
Pixverse's latest V6 Model
Pixverse logo
pixverse/v6/image-to-video

Pixverse's latest V6 Model

image-to-video
Kling 2.5 Turbo Pro: Top-tier text-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
Kling logoKling logo
kling-video/v2.5-turbo/pro/text-to-video

Kling 2.5 Turbo Pro: Top-tier text-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

animation
stylized
text-to-video
Kling 2.1 Pro is an advanced endpoint for the Kling 2.1 model, offering professional-grade videos with enhanced visual fidelity, precise camera movements, and dynamic motion control, perfect for cinematic storytelling.
Kling logoKling logo
kling-video/v2.1/pro/image-to-video

Kling 2.1 Pro is an advanced endpoint for the Kling 2.1 model, offering professional-grade videos with enhanced visual fidelity, precise camera movements, and dynamic motion control, perfect for cinematic storytelling.

deprecated
image-to-video
Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind
Google logoGoogle logo
veo3.1/image-to-video

Veo 3.1 is the latest state-of-the art video generation model from Google DeepMind

image-to-video

Best Avatar Models

Top models for generating talking avatars, lip-sync videos, and expressive character performances.

Generate high fidelity, studio quality videos of your avatar speaking or singing using the Aurora from Creatify team!
creatify/aurora

Generate high fidelity, studio quality videos of your avatar speaking or singing using the Aurora from Creatify team!

lipsync
image-to-video
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
Veed logoVeed logo
veed/fabric-1.0

VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video

lipsync
avatar
image-to-video
The Avatar X API offers access to Mirage's most advanced generation model yet, delivering
industry-leading identity preservation and expressivity in AI video
mirage-api/avatar-x/reference-to-video

The Avatar X API offers access to Mirage's most advanced generation model yet, delivering industry-leading identity preservation and expressivity in AI video

avatar
lipsync
talking-head
image-to-video
Create natural HeyGen Avatar V digital twin videos from text or audio, with lip-sync, optional backgrounds, captions, and MP4/WebM output.
Heygen logo
heygen/avatar5/digital-twin

Create natural HeyGen Avatar V digital twin videos from text or audio, with lip-sync, optional backgrounds, captions, and MP4/WebM output.

avatar
digital-twin
talking-avatar
text-to-video
sync-3 most powerful lipsync model yet, featuring native visual intelligence for professional-quality video.
sync-lipsync/v3

sync-3 most powerful lipsync model yet, featuring native visual intelligence for professional-quality video.

stylized
transform
lipsync
video-to-video
Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
Bytedance logoBytedance logo
bytedance/omnihuman/v1.5

Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.

lipsync
image-to-video
MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.
ai-avatar/single-text

MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.

stylized
transform
image-to-video
Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization with Sync Lipsync 2.0 model
sync-lipsync/v2

Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization with Sync Lipsync 2.0 model

animation
lip sync
video-to-video

Audio Models

Models for speech, music, sound effects, and audio generation across a wide range of use cases.

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.
sonilo/v1.1/text-to-music

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

stylized
transform
lipsync
text-to-audio
Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
sonilo/v1.1/text-to-sound-effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

sfx
audio
effects
text-to-audio
Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.
sonilo/v1.1/video-to-sound-effects

Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.

sfx
audio
effects
video-to-audio
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
inworld-tts

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

inworld
tts
text-to-speech
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
chatterbox/text-to-speech

Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.

text-to-speech
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
Minimax logo
minimax/speech-02-hd

Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

speech
text-to-speech
Clone dialog voices from a sample audio and generate dialogs from text prompts using the Dia TTS which leverages advanced AI techniques to create high-quality text-to-speech.
dia-tts/voice-clone

Clone dialog voices from a sample audio and generate dialogs from text prompts using the Dia TTS which leverages advanced AI techniques to create high-quality text-to-speech.

speech
audio-to-audio
Generate synced sounds for any video, and return the new sound track (like MMAudio)
mirelo-ai/sfx-v1/video-to-audio

Generate synced sounds for any video, and return the new sound track (like MMAudio)

sfx
video-to-audio

Text to Music APIs

Everything you need to start making music with AI

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.
sonilo/v1.1/text-to-music

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

stylized
transform
lipsync
text-to-audio
Lyria 3 Pro is the latest music model from Google
Google logoGoogle logo
lyria3/pro

Lyria 3 Pro is the latest music model from Google

audio
sfx
text-to-audio
Lyria 3 is most recent music model from Google
Google logoGoogle logo
lyria3

Lyria 3 is most recent music model from Google

audio
music
sfx
text-to-audio
Generate high quality, realistic music with fine controls using Elevenlabs Music!
ElevenLabs logoElevenLabs logo
elevenlabs/music

Generate high quality, realistic music with fine controls using Elevenlabs Music!

music
text-to-music
text-to-audio
MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
Minimax logo
minimax-music/v2.6

MiniMax Music 2.6 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

stylized
transform
lipsync
text-to-audio
MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
Minimax logo
minimax-music/v2.5

MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

stylized
transform
lipsync
text-to-audio
Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
Minimax logo
minimax-music

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
text-to-audio
Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
Minimax logo
minimax-music/v2

Generate music from text prompts using the MiniMax Music 2.0 model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
audio
text-to-audio

Best Lora Trainers

Training endpoints for creating and fine-tuning custom LoRA models for personalization and style adaptation.

FLUX LoRA training optimized for portrait generation, with bright highlights, excellent prompt following and highly detailed results.
Black Forest Labs logoBlack Forest Labs logo
flux-lora-portrait-trainer

FLUX LoRA training optimized for portrait generation, with bright highlights, excellent prompt following and highly detailed results.

lora
personalization
training
LoRA trainer for FLUX.1 Kontext [dev]
Black Forest Labs logoBlack Forest Labs logo
flux-kontext-trainer

LoRA trainer for FLUX.1 Kontext [dev]

training
Train styles, people and other subjects at blazing speeds.
Black Forest Labs logoBlack Forest Labs logo
flux-lora-fast-training

Train styles, people and other subjects at blazing speeds.

lora
personalization
training
Train custom LoRAs for Wan-2.1 T2V 14B
Alibaba logoAlibaba logo
wan-trainer/t2v-14b

Train custom LoRAs for Wan-2.1 T2V 14B

lora
training
Qwen Image LoRA training
Alibaba logoAlibaba logo
qwen-image-trainer

Qwen Image LoRA training

deprecated
lora
personalization
training
Fine-tune FLUX.2 [klein] 4B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
Black Forest Labs logoBlack Forest Labs logo
flux-2-klein-4b-base-trainer

Fine-tune FLUX.2 [klein] 4B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.

deprecated
training
Fine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.
Black Forest Labs logoBlack Forest Labs logo
flux-2-klein-9b-base-trainer

Fine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.

training
Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
Black Forest Labs logoBlack Forest Labs logo
flux-2-trainer-v2

Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.

training

Virtual Try On APIs

Virtually try on different outfits and character styles with our collection of APIs.

Bria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.
new
Bria AI logoBria AI logo
bria/fibo-edit-1.5/virtual-try-on

Bria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.

virtual try-on
fashion tech
apparel
image-to-image
Realtime Try On experience with Decart Lucy 2.1 VTON
Decart logo
decart/lucy2-vton/realtime

Realtime Try On experience with Decart Lucy 2.1 VTON

video-to-video
Try on clothes virtually by combining person and clothing images.
image-apps-v2/virtual-try-on

Try on clothes virtually by combining person and clothing images.

fashion
try-on
virtual-try-on
image-to-image
Kling Kolors Virtual TryOn v1.5 is a high quality image based Try-On endpoint which can be used for commercial try on.
Kling logoKling logo
kling/v1-5/kolors-virtual-try-on

Kling Kolors Virtual TryOn v1.5 is a high quality image based Try-On endpoint which can be used for commercial try on.

deprecated
try-on
fashion
clothing
image-to-image
FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.
fashn/tryon/v1.6

FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.

try-on
fashion
clothing
image-to-image
FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.
fashn/tryon/v1.5

FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.

try-on
fashion
clothing
image-to-image
Virtual clothing try-on (2 images: person + garment)
Black Forest Labs logoBlack Forest Labs logo
flux-2-lora-gallery/virtual-tryon

Virtual clothing try-on (2 images: person + garment)

stylized
transform
image-to-image
Leffa Virtual TryOn is a high quality image based Try-On endpoint which can be used for commercial try on.
leffa/virtual-tryon

Leffa Virtual TryOn is a high quality image based Try-On endpoint which can be used for commercial try on.

try-on
fashion
clothing
image-to-image

Image to 3D Model APIs

Run the best image-to-3D models on fal

Generate 3D models from images with Hunyuan 3D Pro
hunyuan-3d/v3.1/pro/image-to-3d

Generate 3D models from images with Hunyuan 3D Pro

3d
hunyuan
image-to-3d
Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.
trellis-2

Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.

image-to-3d
image-to-3d
Generate high-quality 3D models from a single image using Tripo H3.1.
tripo3d/h3.1/image-to-3d

Generate high-quality 3D models from a single image using Tripo H3.1.

3d
3d-generation
tripo
image-to-3d
Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.
trellis

Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-3d
Transform your photos into ultra-high-resolution 3D models in seconds. Film-quality geometry with PBR textures, ready for games, e-commerce, and 3D printing.
hunyuan3d-v3/image-to-3d

Transform your photos into ultra-high-resolution 3D models in seconds. Film-quality geometry with PBR textures, ready for games, e-commerce, and 3D printing.

image-to-3d
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
hyper3d/rodin/v2.5

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

image-to-3d
Meshy 7.1 generates 3D models from a single image, with standard, low-poly, and Smart Topology modes, optional textures and PBR maps, and geometry resolution up to 4K.
new
meshy/v7.1/image-to-3d

Meshy 7.1 generates 3D models from a single image, with standard, low-poly, and Smart Topology modes, optional textures and PBR maps, and geometry resolution up to 4K.

3d
textures
pbr
image-to-3d
Generate 3D models from multiple view images using Tripo H3.1.
tripo3d/h3.1/multiview-to-3d

Generate 3D models from multiple view images using Tripo H3.1.

3d
multiview-to-3d
3d-generation
image-to-3d

Text to 3D Model APIs

This is our collection of the best text-to-3D model APIs available on fal.

Generate 3D models from text descriptions using Tripo H3.1.
tripo3d/h3.1/text-to-3d

Generate 3D models from text descriptions using Tripo H3.1.

3d
3d-generation
tripo
text-to-3d
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
hunyuan-motion

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

motion
text-to-3d
Generate 3D models from text prompts with Hunyuan 3D Pro
hunyuan-3d/v3.1/pro/text-to-3d

Generate 3D models from text prompts with Hunyuan 3D Pro

3d
hunyuan
text-to-3d
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
meshy/v6/text-to-3d

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

text-to-3d
Meshy 7.1 generates 3D models from text prompts, with untextured preview and textured full modes, standard, low-poly, and Smart Topology options, and geometry resolution up to 4K.
new
meshy/v7.1/text-to-3d

Meshy 7.1 generates 3D models from text prompts, with untextured preview and textured full modes, standard, low-poly, and Smart Topology options, and geometry resolution up to 4K.

3d
textures
pbr
text-to-3d
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
hyper3d/rodin/v2.5/text-to-3d

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

text-to-3d
Tripo P2 generates 3D models from a text prompt, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.
new
tripo3d/p2/text-to-3d

Tripo P2 generates 3D models from a text prompt, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.

3d
3d-generation
low-poly
text-to-3d
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
hyper3d/rodin/v2.5/text-to-3d/fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

text-to-3d

Best Utility Models

Specialized models for supporting tasks like background removal, nsfw detection, upscaling and much more.

Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.
Bria AI logoBria AI logo
bria/video/background-removal/v3

Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.

video-to-video
Professional generative image upscaling powered by Topaz Labs. Wonder 3.5 leads the range, with Redefine for prompt-guided detail and Recovery for extreme low-resolution sources. Best for rebuilding sharp detail in small or blurry images.
Topaz Labs logo
topaz/upscale/image/generative

Professional generative image upscaling powered by Topaz Labs. Wonder 3.5 leads the range, with Redefine for prompt-guided detail and Recovery for extreme low-resolution sources. Best for rebuilding sharp detail in small or blurry images.

upscale
image
image-to-image
Professional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.
Topaz Labs logo
topaz/upscale/video/generative

Professional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.

upscale
video
video-to-video
Automatically remove backgrounds from videos -perfect for creating clean, professional content without a green screen.
Bria AI logoBria AI logo
bria/video/background-removal

Automatically remove backgrounds from videos -perfect for creating clean, professional content without a green screen.

background-removal
video-to-video
Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04
Bria AI logoBria AI logo
bria/background/remove

Bria RMBG 2.0 enables seamless removal of backgrounds from images, ideal for professional editing tasks. Trained exclusively on licensed data for safe and risk-free commercial use. Model weights for commercial use are available here: https://share-eu1.hsforms.com/2GLpEVQqJTI2Lj7AMYwgfIwf4e04

background removal
image segmentation
high resolution
image-to-image
Remove video backgrounds in real time with Bria’s VRMBG 3.0 model. Built for live streaming, real-time video apps, content creation, and low-latency workflows that need fast, accurate background removal.
Bria AI logoBria AI logo
bria/video/background-removal/realtime

Remove video backgrounds in real time with Bria’s VRMBG 3.0 model. Built for live streaming, real-time video apps, content creation, and low-latency workflows that need fast, accurate background removal.

bria
video
background-removal
video-to-video
Professional SDR-to-HDR conversion powered by Topaz Labs. Hyperion 2.5 redistributes luminance and color while preserving detail in text, faces and motion. Best for giving flat SDR footage a true HDR look.
Topaz Labs logo
topaz/sdr-to-hdr/video

Professional SDR-to-HDR conversion powered by Topaz Labs. Hyperion 2.5 redistributes luminance and color while preserving detail in text, faces and motion. Best for giving flat SDR footage a true HDR look.

sdr
video
hdr
video-to-video
Predict whether an image is NSFW or SFW.
x-ailab/nsfw

Predict whether an image is NSFW or SFW.

filter
safety
utility
vision

Text To Image APIs

Use the latest state of the art text to image model APIs

Nano Banana Pro is Google's new state-of-the-art image generation and editing model
Google logoGoogle logo
nano-banana-pro

Nano Banana Pro is Google's new state-of-the-art image generation and editing model

realism
typography
text-to-image
FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/schnell

FLUX.1 [schnell] is a 12 billion parameter flow transformer that generates high-quality images from text in 1 to 4 steps, suitable for personal and commercial use.

text-to-image
Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model
Google logoGoogle logo
nano-banana-2

Nano Banana 2 is Google's new state-of-the-art fast image generation and editing model

text-to-image
GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.
OpenAI logoOpenAI logo
openai/gpt-image-2

GPT Image 2, OpenAI's latest image model, is capable of creating extremely detailed images with fine typography.

gpt-image-2
openai
typography
text-to-image
OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/flare/text-to-image

OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.

realism
typography
stylized
text-to-image
FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.
Black Forest Labs logoBlack Forest Labs logo
flux/dev

FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.

text-to-image
OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.
OpenAI logoOpenAI logo
openai/gpt-image-2.5/sunburst/text-to-image

OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.

realism
typography
stylized
text-to-image
Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows
Black Forest Labs logoBlack Forest Labs logo
flux-2-pro

Image editing with FLUX.2 [pro] from Black Forest Labs. Ideal for high-quality image manipulation, style transfer, and sequential editing workflows

text-to-image