Wan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications
Alibaba logo
wan/v2.2-14b/speech-to-video

Wan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications

talking-head
audio-to-video
Generate videos with a single prompt. Describe what you want in plain text, and the agent handles avatar selection, scripting, scene composition - all in one.
Heygen logo
heygen/v3/video-agent

Generate videos with a single prompt. Describe what you want in plain text, and the agent handles avatar selection, scripting, scene composition - all in one.

text-to-video
Generate video with audio from images using LTX-2.3 Distilled
LTX logo
ltx-2.3-22b/distilled/image-to-video

Generate video with audio from images using LTX-2.3 Distilled

image-to-video
Restyle videos up to 30 min long - maintaining maximum detail quality.
Decart logo
decart/lucy-restyle

Restyle videos up to 30 min long - maintaining maximum detail quality.

video-edit
video-to-video
Modify a portion of provided audio with lyrics and/or style using ACE-Step
ace-step/audio-inpaint

Modify a portion of provided audio with lyrics and/or style using ACE-Step

audio-inpaint
audio-repaint
audio-to-audio
Hunyuan Video is an Open video generation model with high visual quality, motion diversity, text-video alignment, and generation stability. This endpoint generates videos from text descriptions.
hunyuan-video

Hunyuan Video is an Open video generation model with high visual quality, motion diversity, text-video alignment, and generation stability. This endpoint generates videos from text descriptions.

motion
text-to-video
Generate video clips from your images using MiniMax Video model
Minimax logo
minimax/video-01-live/image-to-video

Generate video clips from your images using MiniMax Video model

motion
transformation
image-to-video
Heygen Translate Model with Extreme Precision
Heygen logo
heygen/v2/translate/precision

Heygen Translate Model with Extreme Precision

video-to-video
ImagineArt 2.0 is ImagineArt's latest state-of-the-art visual reasoning text-to-image model, generating high-fidelity, professional-grade visuals with lifelike realism, cinematic effects, and strong aesthetic quality.
imagineart/imagineart-2.0-preview/text-to-image

ImagineArt 2.0 is ImagineArt's latest state-of-the-art visual reasoning text-to-image model, generating high-fidelity, professional-grade visuals with lifelike realism, cinematic effects, and strong aesthetic quality.

stylized
transform
typography
text-to-image
Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.
Minimax logo
minimax-music/v1.5

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

music
text-to-audio
DiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.
diffrhythm

DiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.

music
text-to-audio
Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery
pixelcut/product-photo

Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery

utility
editing
image-to-image
Replace backgrounds existing images with Ideogram V3's replace background feature. Create variations and adaptations while preserving core elements and adding new creative directions through prompt guidance.
Ideogram logo
ideogram/v3/replace-background

Replace backgrounds existing images with Ideogram V3's replace background feature. Create variations and adaptations while preserving core elements and adding new creative directions through prompt guidance.

image-to-image
Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.
recraft/v4/pro/text-to-vector

Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.

text-to-vector
text-to-image
Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
Kling logo
kling-video/o3/4k/text-to-video

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

stylized
transform
lipsync
text-to-video
Removes objects and their visual effects using natural language, replacing them with contextually appropriate content
object-removal

Removes objects and their visual effects using natural language, replacing them with contextually appropriate content

utility
editing
image-to-image
Interpolate images with FILM - Frame Interpolation for Large Motion
film

Interpolate images with FILM - Frame Interpolation for Large Motion

interpolation
image-to-image
Image-to-image editing with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.
Black Forest Labs logo
flux-2/klein/4b/base/edit/lora

Image-to-image editing with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.

image-to-image
FLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.
Black Forest Labs logo
flux-control-lora-depth/image-to-image

FLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.

lora
style transfer
image-to-image
Do high precision video upscaling that respects the original video perfectly using Crystal Upscaler's new video upscaling method!
clarityai/crystal-video-upscaler

Do high precision video upscaling that respects the original video perfectly using Crystal Upscaler's new video upscaling method!

upscale
video-to-video
Generate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create high-quality text-to-speech.
Kling logo
kling-video/v1/tts

Generate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create high-quality text-to-speech.

audio
text-to-speech
Run Any Stable Diffusion model with customizable LoRA weights.
lora/image-to-image

Run Any Stable Diffusion model with customizable LoRA weights.

diffusion
lora
customization
image-to-image
Generate video with synchronized audio from a text prompt using MiniMax H3; load a trained LoRA at adjustable strength to lock in style, character, or motion.
Minimax logo
minimax/h3/text-to-video/lora

Generate video with synchronized audio from a text prompt using MiniMax H3; load a trained LoRA at adjustable strength to lock in style, character, or motion.

utility
editing
text-to-video
Automatically generates text captions for your videos from the audio as per text colour/font specifications
auto-caption

Automatically generates text captions for your videos from the audio as per text colour/font specifications

captioning
video
video-to-video
Showing 721 to 744 of 1493 results