VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
Veed logo
veed/fabric-1.0

VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video

lipsync
avatar
image-to-video
Google's famous original image generation and editing model, a.k.a Nano Banana
Google logo
gemini-25-flash-image

Google's famous original image generation and editing model, a.k.a Nano Banana

text-to-image
Transfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
Kling logo
kling-video/v3/standard/motion-control

Transfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.

stylized
transform
editing
video-to-video
Wan-2.2 Turbo image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
Alibaba logo
wan/v2.2-a14b/image-to-video/turbo

Wan-2.2 Turbo image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.

image-to-video
Bria Eraser enables precise removal of unwanted objects from images while maintaining high-quality outputs. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us
Bria AI logo
bria/eraser

Bria Eraser enables precise removal of unwanted objects from images while maintaining high-quality outputs. Trained exclusively on licensed data for safe and risk-free commercial use. Access the model's source code and weights: https://bria.ai/contact-us

image editing
object removal
image-to-image
ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.
Bytedance logo
bytedance/seedance-2.0/fast/text-to-video

ByteDance's most advanced text-to-video model, fast tier. Lower latency and cost with cinematic output, native audio, multi-shot editing, and director-level camera control.

stylized
transform
lipsync
text-to-video
Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs.
Ideogram logo
ideogram/v4

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs.

realism
typography
stylized
text-to-image
Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.
Krea logo
krea-2/turbo

Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.

stylized
transform
typography
text-to-image
Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes
Alibaba logo
alibaba/qwen-image-3/edit

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

stylized
transform
typography
image-to-image
Open source text-to-audio model.
stable-audio

Open source text-to-audio model.

music
text-to-audio
Pixverse's latest V6 Model
Pixverse logo
pixverse/v6/image-to-video

Pixverse's latest V6 Model

image-to-video
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
Bytedance logo
bytedance/seedance-2.0/mini/image-to-video

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

stylized
transform
lipsync
image-to-video
Upscale videos with Bytedance's video upscaler.
Bytedance logo
bytedance-upscaler/upscale/video

Upscale videos with Bytedance's video upscaler.

upscaler
video
bytedance
video-to-video
FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.
Black Forest Labs logo
flux-2-max

FLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.

flux2
max
text-to-image
Merge videos with standalone audio files or audio from video files.
ffmpeg-api/merge-audio-video

Merge videos with standalone audio files or audio from video files.

ffmpeg
video-to-video
Kling AI Avatar v2 Pro: The premium endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters
Kling logo
kling-video/ai-avatar/v2/pro

Kling AI Avatar v2 Pro: The premium endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters

image-to-video
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.
new
Google logo
google/gemini-omni-flash/v1.1/edit

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.

utility
editing
transform
video-to-video
LatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.
latentsync

LatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.

animation
lip sync
video-to-video
Kling LipSync is an audio-to-video model that generates realistic lip movements from audio input.
Kling logo
kling-video/lipsync/audio-to-video

Kling LipSync is an audio-to-video model that generates realistic lip movements from audio input.

audio to video
lipsync
text-to-video
A audio understanding model to analyze audio content and answer questions about what's happening in the audio based on user prompts.
audio-understanding

A audio understanding model to analyze audio content and answer questions about what's happening in the audio based on user prompts.

utility
audio
audio-to-audio
Wan 2.5 image-to-video model.
Alibaba logo
wan-25-preview/image-to-video

Wan 2.5 image-to-video model.

image-to-video
Generate videos with audio from text using Grok Imagine Video.
xAI logo
xai/grok-imagine-video/text-to-video

Generate videos with audio from text using Grok Imagine Video.

xai
grok
t2v
text-to-video
Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model
Google logo
gemini-3-pro-image-preview

Gemini 3 Pro Image (a.k.a Nano Banana Pro) is Google's state-of-the-art high-fidelity image generation and editing model

realism
typography
text-to-image
fal-ai/wan/v2.2-A14B/image-to-video
Alibaba logo
wan/v2.2-a14b/image-to-video

fal-ai/wan/v2.2-A14B/image-to-video

image-to-video
Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.
trellis

Generate 3D models from your images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-3d
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
Alibaba logo
alibaba/wan-3.0/text-to-video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

stylized
transform
lipsync
text-to-video
Generate text from speech using ElevenLabs advanced speech-to-text model.
ElevenLabs logo
elevenlabs/speech-to-text

Generate text from speech using ElevenLabs advanced speech-to-text model.

speech
speech-to-text
An endpoint for personalized image generation using Flux as per given description.
Black Forest Labs logo
flux-pulid

An endpoint for personalized image generation using Flux as per given description.

personalization
style transfer
image-to-image
Showing 169 to 196 of 1504 results