
Generate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA

Luma Uni-1 Max Edit applies text-guided edits to a source image at maximum fidelity, holding the original structure while honoring reference images for precise, high-detail revisions.

Edit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.

US hosted version of ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.

Generate long videos from prompts and images using LTX Video-0.9.8 13B Distilled and custom LoRA

Image2SVG transforms raster images into clean vector graphics, preserving visual quality while enabling scalable, customizable SVG outputs with precise control over detail levels.

Image editing with HY-WU. Transfer outfits, swap faces, and blend textures instantly—no finetuning needed, just describe what you want and provide reference images.

Pixverse Effects

Removes mask-selected objects and their visual effects, seamlessly reconstructing the scene with contextually appropriate content.

Happy Horse 1.1 is Alibaba's #1-ranked video model. This text-to-video endpoint generates 1080p video with synchronized native audio and multilingual lip-sync from a text prompt alone.

Recraft V4.1 Pro Vector generates large-format, fully editable SVGs with the structural clarity professional illustrators expect. Built for poster art, complex brand assets, and detailed scene illustration, it scales without losing geometric integrity.

Wan-2.1 flf2v generates dynamic videos by intelligently bridging a given first frame to a desired end frame through smooth, coherent motion sequences.

Wan 2.2's 5B model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

Generate long, expressive multi-voice speech using Microsoft's powerful TTS

Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.

HiDream-I1 fast is a new open-source image generative foundation model with 17B parameters that achieves state-of-the-art image generation quality within 16 steps.

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.

Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.

Moondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.

Unified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.
![Realtime generation with FLUX.2 [klein] from Black Forest Labs.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a8d5092%2FvaTm5if3zW-sNx3VgjI2T_2b97424cac3e4f62bebb30ddf1aa1d4b.jpg/tr:w-1920,q-80/vaTm5if3zW-sNx3VgjI2T_2b97424cac3e4f62bebb30ddf1aa1d4b.webp)
Realtime generation with FLUX.2 [klein] from Black Forest Labs.

Remove all text and writing from images while preserving the background and natural appearance.