An endpoint for re-lighting photos and changing their backgrounds per a given description
iclight-v2

An endpoint for re-lighting photos and changing their backgrounds per a given description

relighting
editing
image-to-image
Bria Background Replace allows for efficient swapping of backgrounds in images via text prompts or reference image, delivering realistic and polished results. Trained exclusively on licensed data for safe and risk-free commercial use
Bria AI logo
bria/background/replace

Bria Background Replace allows for efficient swapping of backgrounds in images via text prompts or reference image, delivering realistic and polished results. Trained exclusively on licensed data for safe and risk-free commercial use

image editing
image-to-image
Generate 3D models from text prompts with Hunyuan 3D Pro
hunyuan-3d/v3.1/pro/text-to-3d

Generate 3D models from text prompts with Hunyuan 3D Pro

3d
hunyuan
text-to-3d
A fal.ai endpoint that stitches an ordered list of images into an MP4 video by holding each image for a specified number of frames at a configurable frame rate
ffmpeg-api/images-to-video

A fal.ai endpoint that stitches an ordered list of images into an MP4 video by holding each image for a specified number of frames at a configurable frame rate

utility
editing
image-to-video
MiniMax Hailuo-2.3 Text To Video API (Standard, 768p): Advanced text-to-video generation model with 768p resolution
Minimax logo
minimax/hailuo-2.3/standard/text-to-video

MiniMax Hailuo-2.3 Text To Video API (Standard, 768p): Advanced text-to-video generation model with 768p resolution

text-to-video
Generate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.
Alibaba logo
alibaba/happy-horse/text-to-video

Generate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.

happy-horse
text-to-video
PATINA creates seamless high-resolution normal, roughness, basecolor (albedo), height (displacement) and metalness maps from images
patina

PATINA creates seamless high-resolution normal, roughness, basecolor (albedo), height (displacement) and metalness maps from images

pbr
displacement
metalness
image-to-image
Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
new
Kling logo
kling-video/o3/4k/video-to-video/reference

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

utility
editing
video-to-video
Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.
hunyuan3d/v2

Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-3d
FLUX General Inpainting is a versatile endpoint that enables precise image editing and completion, supporting multiple AI extensions including LoRA, ControlNet, and IP-Adapter for enhanced control over inpainting results and sophisticated image modifications.
Black Forest Labs logo
flux-general/inpainting

FLUX General Inpainting is a versatile endpoint that enables precise image editing and completion, supporting multiple AI extensions including LoRA, ControlNet, and IP-Adapter for enhanced control over inpainting results and sophisticated image modifications.

lora
controlnet
ip-adapter
image-to-image
Transfer motion from a video to characters in an image using Dreamactor v2. Great performance for non-human and multiple characters
Bytedance logo
bytedance/dreamactor/v2

Transfer motion from a video to characters in an image using Dreamactor v2. Great performance for non-human and multiple characters

motion-control
dreamactor
video-to-video
Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.
sonilo/v1.1/video-to-sound-effects

Analyzes a video and generates synchronized, royalty-free sound effects timed to visible actions. Returns the generated sound-effects audio track for commercial use.

sfx
audio
effects
video-to-audio
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
chatterbox/speech-to-speech

Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.

speech-to-speech
Lyria 3 is most recent music model from Google
Google logo
lyria3

Lyria 3 is most recent music model from Google

audio
music
sfx
text-to-audio
Rapidly generate 3D models from images using Hunyuan 3D.
hunyuan-3d/v3.1/rapid/image-to-3d

Rapidly generate 3D models from images using Hunyuan 3D.

3d
hunyuan
image-to-3d
FLUX.1 [dev] Fill is a high-performance endpoint for the FLUX.1 [pro] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux-lora-fill

FLUX.1 [dev] Fill is a high-performance endpoint for the FLUX.1 [pro] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

editing
lora
image-to-image
SAM 3D allows for accurate 3D reconstruction of human body shape and position from a single image.
sam-3/3d-body

SAM 3D allows for accurate 3D reconstruction of human body shape and position from a single image.

3d
human
pose
image-to-3d
SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
sam-3/video

SAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.

segmentation
mask
real-time
video-to-video
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.
Black Forest Labs logo
blackforestlabs/flux-3/image-to-video/draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.

stylized
transform
lipsync
image-to-video
Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs FRACTION OF A SECOND.
Ideogram logo
ideogram/v4/instant

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs FRACTION OF A SECOND.

realism
typography
stylized
text-to-image
Wan-2.2 turbo text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
Alibaba logo
wan/v2.2-a14b/text-to-video/turbo

Wan-2.2 turbo text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.

text to video
motion
text-to-video
Inpainting Endpoint for the Qwen Edit Image editing model.
Alibaba logo
qwen-image-edit/inpaint

Inpainting Endpoint for the Qwen Edit Image editing model.

inpainting
qwen-image
image-to-image
Image-to-image editing with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.
Black Forest Labs logo
flux-2/klein/9b/base/edit/lora

Image-to-image editing with LoRA support for FLUX.2 [klein] 9B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.

image-to-image
Extend existing images with Ideogram V3's reframe feature. Create expanded versions and adaptations while preserving main image and adding new creative directions through prompt guidance.
Ideogram logo
ideogram/v3/reframe

Extend existing images with Ideogram V3's reframe feature. Create expanded versions and adaptations while preserving main image and adding new creative directions through prompt guidance.

realism
typography
image-to-image
Kling V3: Latest Kling Image model
Kling logo
kling-image/v3/text-to-image

Kling V3: Latest Kling Image model

text-to-image
SAM 2 is a model for segmenting images automatically. It can return individual masks or a single mask for the entire image.
sam2/auto-segment

SAM 2 is a model for segmenting images automatically. It can return individual masks or a single mask for the entire image.

segmentation
mask
image-to-image
Get EBU R128 loudness normalization from audio files using FFmpeg API.
ffmpeg-api/loudnorm

Get EBU R128 loudness normalization from audio files using FFmpeg API.

ffmpeg
json
Moondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
moondream3-preview/detect

Moondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.

vision
vision
Showing 421 to 448 of 1504 results