Kling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
Kling logo
kling-video/o1/video-to-video/reference

Kling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

video-to-video
FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.
fashn/tryon/v1.5

FASHN v1.5 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 576x864 resolution from both on-model and flat-lay photo references.

try-on
fashion
clothing
image-to-image
Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
stable-audio-3/medium/audio-to-audio

Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.

music
style-transfer
remix
audio-to-audio
Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.
Black Forest Labs logo
flux-2-trainer-v2

Fine-tune FLUX.2 [dev] from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific styles and domains.

training
Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.
Luma AI logo
luma/agent/ray/v3.2/text-to-video

Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.

stylized
transform
lipsync
text-to-video
Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.
Black Forest Labs logo
flux-2/klein/9b/base

Text-to-image generation with FLUX.2 [klein] 9B Base from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities.

text-to-image
FFMPEG Utility for Audio Compression
workflow-utilities/audio-compressor

FFMPEG Utility for Audio Compression

audio-to-audio
SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
sam-3-1/video

SAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.

segmentation
mask
real-time
video-to-video
Generate images from text and images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
Alibaba logo
z-image/turbo/image-to-image/lora

Generate images from text and images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

turbo
z-image
fast
image-to-image
Kling AI Avatar Standard:  Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters
Kling logo
kling-video/v1/standard/ai-avatar

Kling AI Avatar Standard: Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters

stylized
transform
image-to-video
Rembg-enhance is optimized for 2D vector images, 3D graphics, and photos by leveraging matting technology.
smoretalk-ai/rembg-enhance

Rembg-enhance is optimized for 2D vector images, 3D graphics, and photos by leveraging matting technology.

background removal
image editing
utility
image-to-image
LTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates synchronized video and audio from a text prompt in a single pass, in a quality-optimized mode for final, high-fidelity output.
LTX logo
lightricks/ltx-2.5/text-to-video/pro

LTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates synchronized video and audio from a text prompt in a single pass, in a quality-optimized mode for final, high-fidelity output.

stylized
transform
lipsync
text-to-video
MoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.
moondream-next

MoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.

multimodal
vision
Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2
Alibaba logo
wan/v2.2-a14b/image-to-video/lora

Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts and images. This endpoint supports LoRAs made for Wan 2.2

motion
lora
image-to-video
Generates same object from different angles (azimuth/elevation)
Black Forest Labs logo
flux-2-lora-gallery/multiple-angles

Generates same object from different angles (azimuth/elevation)

stylized
transform
image-to-image
Create depth maps using Marigold depth estimation.
imageutils/marigold-depth

Create depth maps using Marigold depth estimation.

depth
utility
image-to-image
Professional-grade creative upscaler that doubles resolution up to 10MP, regenerating sharper textures, refined details, and cleaner faces. Trained exclusively on licensed data for risk-free commercial use.
Bria AI logo
bria/upscale/creative

Professional-grade creative upscaler that doubles resolution up to 10MP, regenerating sharper textures, refined details, and cleaner faces. Trained exclusively on licensed data for risk-free commercial use.

bria
aesthetics
upscaler
image-to-image
MiniMax Hailuo-2.3 Text To Video API (Pro, 1080p): Advanced text-to-video generation model with 1080p resolution
Minimax logo
minimax/hailuo-2.3/pro/text-to-video

MiniMax Hailuo-2.3 Text To Video API (Pro, 1080p): Advanced text-to-video generation model with 1080p resolution

text-to-video
Ideogram Layerize takes an existing flat graphic, removes text, and returns structured text containers you can edit/recompose in html or json format.
Ideogram logo
ideogram/v3/layerize-text

Ideogram Layerize takes an existing flat graphic, removes text, and returns structured text containers you can edit/recompose in html or json format.

stylized
transform
typography
image-to-image
Generates perfectly synced music for any video. Return a licensed music soundtrack ready for commercial use (optional preservation of the original speech in video)
sonilo/v1.1/video-to-video-music

Generates perfectly synced music for any video. Return a licensed music soundtrack ready for commercial use (optional preservation of the original speech in video)

music
editing
restoration
video-to-video
LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
LTX logo
ltx-2.3/audio-to-video

LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.

stylized
transform
lipsync
audio-to-video
Instruct version of Hunyuan-Image 3.0, with internal reasoning capabilities.
hunyuan-image/v3/instruct/text-to-image

Instruct version of Hunyuan-Image 3.0, with internal reasoning capabilities.

hunyuan-image
v3
instruct
text-to-image
Meshy-5 remesh allows you to remesh and export existing 3D models into various formats
meshy/v5/remesh

Meshy-5 remesh allows you to remesh and export existing 3D models into various formats

3d-to-3d
FLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
Black Forest Labs logo
flux-1/krea

FLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

text-to-image
Showing 601 to 624 of 1491 results