
MoonDreamNext Detection is a multimodal vision-language model for gaze detection, bbox detection, point detection, and more.

Fast LoRA trainer for Z-Image-Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

Run SDXL at the speed of light

Get waveform data from audio files using FFmpeg API.

Generate high quality video clips from text and image prompts using PixVerse v4.5

Generate seamlessly tiling photorealistic images from text using Z-Image Turbo

Default parameters with automated optimizations and quality improvements.

Generate video with audio from images using LTX-2.3

Video reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts video plus a prompt and returns text.

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.

VACE Fun for Wan 2.2 A14B from Alibaba-PAI

Ideogram Upscale enhances the resolution of the reference image by up to 2X and might enhance the reference image too. Optionally refine outputs with a prompt for guided improvements.

Transform your 3D video render into realistic using first frame with Ltx 2.3

Generate images from text and edge, depth or pose images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Experiment with different hairstyles, from bald to any style you can imagine, while maintaining natural lighting and realistic results.

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.

Transfer expression from a video to a portrait.

Generate long videos in 720p/30fps from images using LongCat Video Distilled

Post Processing is an endpoint that can enhance images using a variety of techniques including grain, blur, sharpen, and more.

Line art preprocessor.

Run any VLM (Video Language Model) with fal, powered by OpenRouter.

Vidu's latest Q3 Reference to Video Mix model

High-quality text-to-image model by Baidu. Supports English, Chinese, and Japanese prompts with built-in prompt expansion.