
Virtual clothing try-on (2 images: person + garment)
![FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffal_cdn%2Ffal%2FUpscale-2.jpg/tr:w-1920,q-80/Upscale-2.webp)
FLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.

Erase unwanted objects, people, or elements from video with a text prompt. High-fidelity output with strong temporal consistency, trained on licensed data for safe commercial use.

Generate video with audio from text using LTX-2

LongCat image Edit is a 6B parameter image editing model excelling at multilingual text rendering, photorealism and deployment efficiency.

FFMPEG Utility for Blending Videos

Generate high quality video clips from text and image prompts using PixVerse v4
![Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.](https://refinery.fal.media/url/https%3A%2F%2Fv3.fal.media%2Ffiles%2Ftiger%2FfB-RsJ-BW4mrUVAH8oKF2_LOuGVDgg07U8OWbOhhMFt_d6ab08c96ab94da8b6d3e979d634af16.jpg/tr:w-1920,q-80/fB-RsJ-BW4mrUVAH8oKF2_LOuGVDgg07U8OWbOhhMFt_d6ab08c96ab94da8b6d3e979d634af16.webp)
Super fast endpoint for the FLUX.1 [dev] model with LoRA support, enabling rapid and high-quality image generation using pre-trained LoRA adaptations for personalization, specific styles, brand identities, and product-specific outputs.

Use the amazing capabilities of hunyuan image 2.1 to generate images that express the feelings of your text.

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Apply sharpening effects with three modes: basic unsharp mask, smart sharpening with edge preservation, and Contrast Adaptive Sharpening (CAS).

Use the latest pixverse v5.6 model to turn your texts and images into amazing videos.

Use the latest Vidu Q2 models which much more better quality and control on your videos.

Ovi can generate videos with audio from image and text inputs.

Transform your photos into vibrant cool cartoons with bold outlines and rich colors.

Edit videos using plain language and Wan VACE

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

Ray2 Flash is a fast video generative model capable of creating realistic visuals with natural, coherent motion.

Apply artistic styles like impressionism, cubism, or surrealism to your images.

Generate professional, eCommerce-ready product shots by replacing backgrounds with realistic lighting and accurate perspective from a simple text prompt. Trained exclusively on licensed data for safe commercial use.

Wan 2.6 reference-to-video model.

HiDream-I1 dev is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within seconds.

Text-to-Image endpoint for Qwen-Image-Max. Qwen Image Max improves upon the Qwen Image Plus series by enhancing the realism and naturalness of images.

NVIDIA's Logically Consistent and Physics-Aware Image Editing Model