
Wan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications

Generate videos with a single prompt. Describe what you want in plain text, and the agent handles avatar selection, scripting, scene composition - all in one.

Generate video with audio from images using LTX-2.3 Distilled

Restyle videos up to 30 min long - maintaining maximum detail quality.

Modify a portion of provided audio with lyrics and/or style using ACE-Step

Hunyuan Video is an Open video generation model with high visual quality, motion diversity, text-video alignment, and generation stability. This endpoint generates videos from text descriptions.

Generate video clips from your images using MiniMax Video model

Heygen Translate Model with Extreme Precision

ImagineArt 2.0 is ImagineArt's latest state-of-the-art visual reasoning text-to-image model, generating high-fidelity, professional-grade visuals with lifelike realism, cinematic effects, and strong aesthetic quality.

Generate music from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality, diverse musical compositions.

DiffRhythm is a blazing fast model for transforming lyrics into full songs. It boasts the capability to generate full songs in less than 30 seconds.

Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery

Replace backgrounds existing images with Ideogram V3's replace background feature. Create variations and adaptations while preserving core elements and adding new creative directions through prompt guidance.

Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

Removes objects and their visual effects using natural language, replacing them with contextually appropriate content

Interpolate images with FILM - Frame Interpolation for Large Motion
![Image-to-image editing with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8b09b3%2Fck_nRVKlUom4-4_5qfG7t_117a3ccf9f9541aeb83e5ffa75564e6d.jpg/tr:w-1920,q-80/ck_nRVKlUom4-4_5qfG7t_117a3ccf9f9541aeb83e5ffa75564e6d.webp)
Image-to-image editing with LoRA support for FLUX.2 [klein] 4B Base from Black Forest Labs. Specialized style transfer and domain-specific modifications.

FLUX Control LoRA Depth is a high-performance endpoint that uses a control image using a depth map to transfer structure to the generated image and another initial image to guide color.

Do high precision video upscaling that respects the original video perfectly using Crystal Upscaler's new video upscaling method!

Generate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create high-quality text-to-speech.

Run Any Stable Diffusion model with customizable LoRA weights.

Generate video with synchronized audio from a text prompt using MiniMax H3; load a trained LoRA at adjustable strength to lock in style, character, or motion.
Automatically generates text captions for your videos from the audio as per text colour/font specifications