MoonDreamNext Detection is a multimodal vision-language model for gaze detection, bbox detection, point detection, and more.
moondream-next/detection

MoonDreamNext Detection is a multimodal vision-language model for gaze detection, bbox detection, point detection, and more.

multimodal
image-to-image
Fast LoRA trainer for Z-Image-Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.
Alibaba logo
z-image-turbo-trainer-v2

Fast LoRA trainer for Z-Image-Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

lora
personalization
trainer
training
Run SDXL at the speed of light
fast-lightning-sdxl/image-to-image

Run SDXL at the speed of light

diffusion
lightning
editing
image-to-image
Get waveform data from audio files using FFmpeg API.
ffmpeg-api/waveform

Get waveform data from audio files using FFmpeg API.

ffmpeg
json
Generate high quality video clips from text and image prompts using PixVerse v4.5
Pixverse logo
pixverse/v4.5/text-to-video

Generate high quality video clips from text and image prompts using PixVerse v4.5

stylized
transform
text-to-video
Generate seamlessly tiling photorealistic images from text using Z-Image Turbo
Alibaba logo
z-image/turbo/tiling

Generate seamlessly tiling photorealistic images from text using Z-Image Turbo

z-image
turbo
seamless
text-to-image
Default parameters with automated optimizations and quality improvements.
fooocus/inpaint

Default parameters with automated optimizations and quality improvements.

stylized
editing
text-to-image
Generate video with audio from images using LTX-2.3
LTX logo
ltx-2.3-22b/image-to-video

Generate video with audio from images using LTX-2.3

image-to-video
Video reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts video plus a prompt and returns text.
nvidia/nemotron-3-nano-omni/video

Video reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts video plus a prompt and returns text.

nemotron
nvidia
video-understanding
video-to-text
Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.
stable-audio-3/small/music/base/audio-to-audio

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.

music
style-transfer
remix
audio-to-audio
VACE Fun for Wan 2.2 A14B from Alibaba-PAI
Alibaba logo
wan-22-vace-fun-a14b/inpainting

VACE Fun for Wan 2.2 A14B from Alibaba-PAI

video-to-video
Ideogram Upscale enhances the resolution of the reference image by up to 2X and might enhance the reference image too. Optionally refine outputs with a prompt for guided improvements.
Ideogram logo
ideogram/upscale

Ideogram Upscale enhances the resolution of the reference image by up to 2X and might enhance the reference image too. Optionally refine outputs with a prompt for guided improvements.

upscaling
high-res
image-to-image
Transform your 3D video render into realistic using first frame with Ltx 2.3
LTX logo
ltx-2.3-quality/render-to-real

Transform your 3D video render into realistic using first frame with Ltx 2.3

3d
video
video-to-video
Generate images from text and edge, depth or pose images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
Alibaba logo
z-image/turbo/controlnet

Generate images from text and edge, depth or pose images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

image-to-image
Experiment with different hairstyles, from bald to any style you can imagine, while maintaining natural lighting and realistic results.
image-editing/hair-change

Experiment with different hairstyles, from bald to any style you can imagine, while maintaining natural lighting and realistic results.

stylized
transform
image-to-image
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
Alibaba logo
wan-vace-14b

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

image-to-video
text-to-video
video-to-video
Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.
new
recraft/v4/style/text-to-vector

Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.

stylized
transform
editing
text-to-image
Transfer expression from a video to a portrait.
live-portrait/image

Transfer expression from a video to a portrait.

expression
animation
image-to-image
Generate long videos in 720p/30fps from images using LongCat Video Distilled
longcat-video/distilled/image-to-video/720p

Generate long videos in 720p/30fps from images using LongCat Video Distilled

image-to-video
Post Processing is an endpoint that can enhance images using a variety of techniques including grain, blur, sharpen, and more.
post-processing

Post Processing is an endpoint that can enhance images using a variety of techniques including grain, blur, sharpen, and more.

stylized
utility
image-to-image
Line art preprocessor.
image-preprocessors/lineart

Line art preprocessor.

preprocess
utility
sketch
image-to-image
Run any VLM (Video Language Model) with fal, powered by OpenRouter.
openrouter/router/video/enterprise

Run any VLM (Video Language Model) with fal, powered by OpenRouter.

video-to-text
Vidu's latest Q3 Reference to Video Mix model
vidu/q3/reference-to-video/mix

Vidu's latest Q3 Reference to Video Mix model

image-to-video
High-quality text-to-image model by Baidu. Supports English, Chinese, and Japanese prompts with built-in prompt expansion.
ernie-image

High-quality text-to-image model by Baidu. Supports English, Chinese, and Japanese prompts with built-in prompt expansion.

realism
chinese
multilingual
text-to-image
Showing 745 to 768 of 1493 results