Sana v1.5 4.8B is a powerful text-to-image model that generates ultra-high quality 4K images with remarkable detail.
sana/v1.5/4.8b

Sana v1.5 4.8B is a powerful text-to-image model that generates ultra-high quality 4K images with remarkable detail.

text to image
4k
high-quality
text-to-image
Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.
infinitalk/video-to-video

Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.

video-to-video
A high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.
kokoro/italian

A high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.

speech
text-to-audio
Generate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
Alibaba logo
z-image/turbo/controlnet/lora

Generate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

turbo
z-image
fast
image-to-image
Finegrain Eraser removes objects—along with their shadows, reflections, and lighting artifacts—using only natural language, seamlessly filling the scene with contextually accurate content.
finegrain-eraser

Finegrain Eraser removes objects—along with their shadows, reflections, and lighting artifacts—using only natural language, seamlessly filling the scene with contextually accurate content.

utility
editing
image-to-image
Edit images with a text prompt using Emu 3.5 Image
emu-3.5-image/edit-image

Edit images with a text prompt using Emu 3.5 Image

image-to-image
Generate video with audio from audio, text and images using LTX-2
LTX logo
ltx-2-19b/audio-to-video

Generate video with audio from audio, text and images using LTX-2

audio-to-video
Generate videos from prompts using LTX Video-0.9.5
LTX logo
ltx-video-v095

Generate videos from prompts using LTX Video-0.9.5

video
text-video
text-to-video
Image-to-image editing with Step1X-Edit v2 from StepFun. Reasoning-enhanced modifications through a thinking–editing–reflection loop with MLLM world knowledge for abstract instruction comprehension.
stepx-edit2

Image-to-image editing with Step1X-Edit v2 from StepFun. Reasoning-enhanced modifications through a thinking–editing–reflection loop with MLLM world knowledge for abstract instruction comprehension.

image-to-image
MiDaS depth estimation preprocessor.
image-preprocessors/midas

MiDaS depth estimation preprocessor.

depth
preprocess
utility
image-to-image
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux-1/schnell/redux

FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

image-to-image
Generate high quality images from text prompts using CogView4. Longer text prompts will result in better quality images.
cogview4

Generate high quality images from text prompts using CogView4. Longer text prompts will result in better quality images.

stylized
text-to-image
FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.
Black Forest Labs logo
blackforestlabs/flux-3/draft-enhance

FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.

stylized
transform
lipsync
video-to-video
A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.
kokoro/mandarin-chinese

A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.

speech
text-to-audio
HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.
hidream-i1-full/image-to-image

HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

hidream
image-to-image
Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.
ovis-image

Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.

ovis-image
artistic
text-to-image
See how you or others might look at different ages, from younger to older, while preserving core facial features.
image-editing/age-progression

See how you or others might look at different ages, from younger to older, while preserving core facial features.

stylized
transform
image-to-image
Generate long videos in 720p/30fps from images using LongCat Video
longcat-video/image-to-video/720p

Generate long videos in 720p/30fps from images using LongCat Video

image-to-video
FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
Black Forest Labs logo
flux-1/srpo

FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

text-to-image
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
Alibaba logo
wan-vace-14b/reframe

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

reframe
video-to-video
Texture an existing geometry mesh using a reference image with Hi3D.
hitem3d/hi3d/texture

Texture an existing geometry mesh using a reference image with Hi3D.

3d-to-3d
Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.
hunyuan3d/v2/mini/turbo

Generate 3D models from your images using Hunyuan 3D. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-3d
Place your subject in any scene you imagine, from enchanted forests to urban settings, with professional composition and lighting
image-editing/scene-composition

Place your subject in any scene you imagine, from enchanted forests to urban settings, with professional composition and lighting

stylized
transform
image-to-image
Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.
krea-wan-14b/video-to-video

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.

video-to-video
Showing 1057 to 1080 of 1496 results