Wan 2.2's 5B model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding
Alibaba logo
wan/v2.2-5b/text-to-video

Wan 2.2's 5B model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

text-to-video
MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
Minimax logo
minimax-music/v2.5

MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

stylized
transform
lipsync
text-to-audio
Generate long, expressive multi-voice speech using Microsoft's powerful TTS
vibevoice/7b

Generate long, expressive multi-voice speech using Microsoft's powerful TTS

multi-speaker
podcast
text-to-speech
Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.
Krea logo
krea-2/turbo/style

Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.

stylized
style transfer
reference image
text-to-image
HiDream-I1 fast is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within 16 steps.
hidream-i1-fast

HiDream-I1 fast is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within 16 steps.

text-to-image
Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.
hidream-o1-image/edit

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

image-to-image
Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation
sadtalker

Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

animation
image-to-video
Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.
mirelo-ai/sfx1.6/text-to-audio

Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.

sfx
text-to-audio
Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.
moondream2/object-detection

Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.

image-to-image
vision
Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.
hidream-o1-image

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

text-to-image
Realtime generation with FLUX.2 [klein] from Black Forest Labs.
Black Forest Labs logo
flux-2/klein/realtime

Realtime generation with FLUX.2 [klein] from Black Forest Labs.

realtime
image-to-image
Remove all text and writing from images while preserving the background and natural appearance.
image-editing/text-removal

Remove all text and writing from images while preserving the background and natural appearance.

stylized
transform
image-to-image
Kling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
Kling logo
kling-video/o1/video-to-video/reference

Kling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

video-to-video
FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.
fashn/tryon/v1.5

FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.

try-on
fashion
clothing
image-to-image
Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
stable-audio-3/medium/audio-to-audio

Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.

music
style-transfer
remix
audio-to-audio
Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
Black Forest Labs logo
flux-2-trainer-v2

Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.

training
Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.
Luma AI logo
luma/agent/ray/v3.2/text-to-video

Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.

stylized
transform
lipsync
text-to-video
Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
Black Forest Labs logo
flux-2/klein/9b/base

Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

text-to-image
FFMPEG Utility for Audio Compression
workflow-utilities/audio-compressor

FFMPEG Utility for Audio Compression

audio-to-audio
SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
sam-3-1/video

SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.

segmentation
mask
real-time
video-to-video
Generate images from text and images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
Alibaba logo
z-image/turbo/image-to-image/lora

Generate images from text and images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

turbo
z-image
fast
image-to-image
Kling AI Avatar Standard:  Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters
Kling logo
kling-video/v1/standard/ai-avatar

Kling AI Avatar Standard: Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters

stylized
transform
image-to-video
Rembg-enhance is optimized for 2D vector images, 3D graphics, and photos by leveraging matting technology.
smoretalk-ai/rembg-enhance

Rembg-enhance is optimized for 2D vector images, 3D graphics, and photos by leveraging matting technology.

background removal
image editing
utility
image-to-image
LTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates synchronized video and audio from a text prompt in a single pass, in a quality-optimized mode for final, high-fidelity output.
LTX logo
lightricks/ltx-2.5/text-to-video/pro

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates synchronized video and audio from a text prompt in a single pass, in a quality-optimized mode for final, high-fidelity output.

stylized
transform
lipsync
text-to-video
MoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.
moondream-next

MoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.

multimodal
vision
Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2
Alibaba logo
wan/v2.2-a14b/image-to-video/lora

Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2

motion
lora
image-to-video
Generates same object from different angles (azimuth/elevation)
Black Forest Labs logo
flux-2-lora-gallery/multiple-angles

Generates same object from different angles (azimuth/elevation)

stylized
transform
image-to-image
Create depth maps using Marigold depth estimation.
imageutils/marigold-depth

Create depth maps using Marigold depth estimation.

depth
utility
image-to-image
Showing 589 to 616 of 1491 results