Edit images with a text prompt using Emu 3.5 Image
emu-3.5-image/edit-image

Edit images with a text prompt using Emu 3.5 Image

image-to-image
A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.
kokoro/hindi

A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.

speech
text-to-audio
Vidu Reference to Video creates videos by using a reference images and combining them with a prompt.
vidu/reference-to-video

Vidu Reference to Video creates videos by using a reference images and combining them with a prompt.

motion
reference
image-to-video
Finegrain Eraser removes any object selected with a bounding box—along with its shadows, reflections, and lighting artifacts—seamlessly reconstructing the scene with contextually accurate content.
finegrain-eraser/bbox

Finegrain Eraser removes any object selected with a bounding box—along with its shadows, reflections, and lighting artifacts—seamlessly reconstructing the scene with contextually accurate content.

utility
editing
image-to-image
Generate long videos from text using LongCat Video Distilled
longcat-video/distilled/text-to-video/480p

Generate long videos from text using LongCat Video Distilled

text-to-video
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux-1/schnell/redux

FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

image-to-image
Kandinsky 5.0 Distilled is a lightweight diffusion model for fast, high-quality text-to-video generation.
kandinsky5/text-to-video/distill

Kandinsky 5.0 Distilled is a lightweight diffusion model for fast, high-quality text-to-video generation.

text-to-video
Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.
Bytedance logo
bernini-r/reference-to-video

Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.

reference
video
stylized
image-to-video
Recraft V3 Create Style is capable of creating unique styles for Recraft V3 based on your images.
recraft/v3/create-style

Recraft V3 Create Style is capable of creating unique styles for Recraft V3 based on your images.

style
vector
personalization
training
A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency.
Bria AI logo
bria/bria_video_eraser/erase/mask

A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency.

bria
erase
video-to-video
Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.
stable-audio-3/medium/base/audio-outpainting

Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.

music
extension
continuation
audio-to-audio
Generate video clips from your prompts using Kling 1.6 (std)
Kling logo
kling-video/v1.6/standard/effects

Generate video clips from your prompts using Kling 1.6 (std)

text-to-video
Lumina-Image-2.0 is a 2 billion parameter flow-based diffusion transforer which features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.
lumina-image/v2

Lumina-Image-2.0 is a 2 billion parameter flow-based diffusion transforer which features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.

diffusion
typography
style
text-to-image
Image generation with BitDance. Fast, high-resolution photorealistic images using an autoregressive LLM— for efficient, high-quality results.
bitdance

Image generation with BitDance. Fast, high-resolution photorealistic images using an autoregressive LLM— for efficient, high-quality results.

text-to-image
Transforms images into comic book style
Black Forest Labs logo
flux-2-lora-gallery/digital-comic-art

Transforms images into comic book style

stylized
transform
text-to-image
Generate HDR from reference video using LTX-2.3
LTX logo
ltx-2.3-quality/hdr

Generate HDR from reference video using LTX-2.3

video-to-video
Interpolate between video frames
amt-interpolation

Interpolate between video frames

interpolation
editing
video-to-video
MiDaS depth estimation preprocessor.
image-preprocessors/midas

MiDaS depth estimation preprocessor.

depth
preprocess
utility
image-to-image
Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.
krea-wan-14b/video-to-video

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.

video-to-video
Vidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image
vidu/image-to-video

Vidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image

motion
image to video
image-to-video
Generate 3D models from multiple view images using Hi3D.
hitem3d/hi3d/multi-view-to-3d

Generate 3D models from multiple view images using Hi3D.

multiview-to-3d
3d
image-to-3d
Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.
new
recraft/v4/create-style

Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.

transform
utility
stylized
training
Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels
sa2va/4b/image

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

multimodal
vision
Generate synced sounds for any video, and return the new sound track (like MMAudio)
mirelo-ai/sfx-v1/video-to-audio

Generate synced sounds for any video, and return the new sound track (like MMAudio)

sfx
video-to-audio
Showing 1081 to 1104 of 1496 results