Generate synced sounds for any video, and return the new sound track (like MMAudio)
mirelo-ai/sfx-v1/video-to-audio

Generate synced sounds for any video, and return the new sound track (like MMAudio)

sfx
video-to-audio
Answer questions from the images.
moondream/batched

Answer questions from the images.

multimodal
vision
A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.
kokoro/hindi

A fast and expressive Hindi text-to-speech model with clear pronunciation and accurate intonation.

speech
text-to-audio
Vidu Q1 Image to Video generates high-quality 1080p videos with exceptional visual quality and motion diversity from a single image
vidu/q1/image-to-video

Vidu Q1 Image to Video generates high-quality 1080p videos with exceptional visual quality and motion diversity from a single image

stylized
transform
image-to-video
Juggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.
rundiffusion-fal/juggernaut-flux/pro/image-to-image

Juggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.

image generation
image-to-image
Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.
new
recraft/v4/create-style

Creates a reusable style from your reference images and returns a style ID you can pass to Recraft V4 Styles image and vector generation.

transform
utility
stylized
training
GOT-OCR2 works on a wide range of tasks, including plain document OCR, scene text OCR, formatted document OCR, and even OCR for tables, charts, mathematical formulas, geometric shapes, molecular formulas and sheet music.
got-ocr/v2

GOT-OCR2 works on a wide range of tasks, including plain document OCR, scene text OCR, formatted document OCR, and even OCR for tables, charts, mathematical formulas, geometric shapes, molecular formulas and sheet music.

optical character recognition
high-res
utility
vision
Edit any image with a natural-language instruction using Bernini-R, changing the weather, materials, objects, or style while preserving the original composition.
Bytedance logo
bernini-r/edit-image

Edit any image with a natural-language instruction using Bernini-R, changing the weather, materials, objects, or style while preserving the original composition.

edit
image
transform
image-to-image
Edit images with a text prompt using Emu 3.5 Image
emu-3.5-image/edit-image

Edit images with a text prompt using Emu 3.5 Image

image-to-image
HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.
hidream-i1-full/image-to-image

HiDream-I1 full is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

hidream
image-to-image
Generate long videos in 720p/30fps from images using LongCat Video
longcat-video/image-to-video/720p

Generate long videos in 720p/30fps from images using LongCat Video

image-to-video
See how you or others might look at different ages, from younger to older, while preserving core facial features.
image-editing/age-progression

See how you or others might look at different ages, from younger to older, while preserving core facial features.

stylized
transform
image-to-image
Generate 3D models from one or more images using ReconViaGen 0.5
reconviagen-0.5

Generate 3D models from one or more images using ReconViaGen 0.5

multi-view
3d-reconstruction
image-to-3d
Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.
krea-wan-14b/video-to-video

Superfast video model based on Wan 2.1 14b by Krea, excelling at real-time video-editing.

video-to-video
Enhance and refine portrait photos with improved clarity and detail.
image-apps-v2/portrait-enhance

Enhance and refine portrait photos with improved clarity and detail.

image-edit
enhancement
image-to-image
Generate high-quality video with audio from text using LTX-2.3
LTX logo
ltx-2.3-quality/text-to-video

Generate high-quality video with audio from text using LTX-2.3

video
text-to-video
FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
Black Forest Labs logo
flux-1/schnell/redux

FLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.

image-to-image
Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.
ovis-image

Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.

ovis-image
artistic
text-to-image
MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.
ai-avatar/single-text

MultiTalk model generates a talking avatar video from an image and text. Converts text to speech automatically, then generates the avatar speaking with lip-sync.

stylized
transform
image-to-video
State-of-the-art open-source model in aesthetic quality
playground-v25/image-to-image

State-of-the-art open-source model in aesthetic quality

artistic
style
image-to-image
Use the latest Vidu Q2 models which much more better quality and control on your videos.
vidu/q2/text-to-video

Use the latest Vidu Q2 models which much more better quality and control on your videos.

text-to-video
Clone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create speeches of yours!
Alibaba logo
qwen-3-tts/clone-voice/0.6b

Clone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create speeches of yours!

clone-voice
voice-clone
audio-to-audio
A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.
kokoro/mandarin-chinese

A highly efficient Mandarin Chinese text-to-speech model that captures natural tones and prosody.

speech
text-to-audio
Transforms images into comic book style
Black Forest Labs logo
flux-2-lora-gallery/digital-comic-art

Transforms images into comic book style

stylized
transform
text-to-image
Showing 1033 to 1056 of 1493 results