
Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library

Add immersive sound effects and background music to your videos using PixVerse sound effects generation

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q using LoRA — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs.

Stable Audio 3 Small Music audio-to-audio is a 459 million parameter latent diffusion model that transforms input music into new variations up to 2 minutes guided by text prompts.

Inpaint images with SD and SDXL

Transform your photos into artistic masterpieces inspired by famous styles like Van Gogh's Starry Night or any artistic style you choose.

LongCat image is a 6B parameter model excelling at multilingual text rendering, photorealism and deployment efficiency.
![Text-to-image generation with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8b09ad%2FV8uFhTiTNXdAgvt1tbJmB_1335a918cf5542539d5954c13b7d0fef.jpg/tr:w-1920,q-80/V8uFhTiTNXdAgvt1tbJmB_1335a918cf5542539d5954c13b7d0fef.webp)
Text-to-image generation with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Custom style adaptation and fine-tuned model variations.

A unified paradigm for audio-video generation

Kling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.

Generate a 3D relief depth map with Hi3D from a single image.

Reimagine and transform your ordinary photos into enchanting Studio Ghibli style artwork

Predict poses from videos.

Framepack is an efficient Image-to-video model that autoregressively generates videos.

Generate high-quality videos with UGC-like avatars from text

DeepSeek Janus-Pro is a novel text-to-image model that unifies multimodal understanding and generation through an autoregressive framework

Generate YouTube thumbnails with custom text

Generate video with audio from audio, text and images using LTX-2 Distilled

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.

Use the latest Vidu Q2 models which much more better quality and control on your videos.

Image-to-image editing with Step1X-Edit v2 from StepFun. Reasoning-enhanced modifications through a thinking–editing–reflection loop with MLLM world knowledge for abstract instruction comprehension.

Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
![FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a9f92d3%2FYI0vlnMufwkKs0eTTXmM7_UVyjAsaK.png/tr:w-1920,q-80/YI0vlnMufwkKs0eTTXmM7_UVyjAsaK.webp)
FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.