
Get waveform data from audio files using FFmpeg API.

Compose videos from multiple media sources using FFmpeg API.

Generate video clips maintaining consistent, realistic facial features and identity across dynamic video content

MoonDreamNext Batch is a multimodal vision-language model for batch captioning.
FLUX1.1 [pro] is an enhanced version of FLUX.1 [pro], improved image generation capabilities, delivering superior composition, detail, and artistic fidelity compared to its predecessor.
![FLUX1.1 [pro] ultra fine-tuned is the newest version of FLUX1.1 [pro] with a fine-tuned LoRA, maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Fflux-pro-11-ultra.webp/tr:w-1920,q-80/flux-pro-11-ultra.webp)
FLUX1.1 [pro] ultra fine-tuned is the newest version of FLUX1.1 [pro] with a fine-tuned LoRA, maintaining professional-grade image quality while delivering up to 2K resolution with improved photo realism.
![FLUX.1 [pro] Fill Fine-tuned is a high-performance endpoint for the FLUX.1 [pro] model with a fine-tuned LoRA that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Ffluxpro.jpg/tr:w-1920,q-80/fluxpro.webp)
FLUX.1 [pro] Fill Fine-tuned is a high-performance endpoint for the FLUX.1 [pro] model with a fine-tuned LoRA that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
![Utilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.](https://refinery.fal.media/url/https%3A%2F%2Fstorage.googleapis.com%2Ffalserverless%2Fgallery%2Fflux_lora.jpg/tr:w-1920,q-80/flux_lora.webp)
Utilize Flux.1 [dev] Controlnet to generate high-quality images with precise control over composition, style, and structure through advanced edge detection and guidance mechanisms.

Hunyuan Video is an Open video generation model with high visual quality, motion diversity, text-video alignment, and generation stability

Generate videos from prompts using CogVideoX-5B

Transform text into stunning videos with TransPixar - an AI model that generates both RGB footage and alpha channels, enabling seamless compositing and creative video effects.

Train Hunyuan Video lora on people, objects, characters and more!

Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization.

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

MoonDreamNext Detection is a multimodal vision-language model for gaze detection, bbox detection, point detection, and more.

MoonDreamNext is a multimodal vision-language model for captioning, gaze detection, bbox detection, point detection, and more.

Generate video clips from your images using Kling 1.6 (std)

Generate video clips from your prompts using Kling 1.6 (std)

Generate video clips from your images using Kling 1.6 (pro)
Automatically generates text captions for your videos from the audio as per text colour/font specifications

Train styles, people and other subjects at blazing speeds.