![Text-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs and custom LoRA.](https://refinery.fal.media/url/https%3A%2F%2Fv3b.fal.media%2Ffiles%2Fb%2F0a928e3b%2FsIC-Ne9BMwZZtBvR3FwKN_9a724704a550471a9df59999e9e1017f.jpg/tr:w-1920,q-80/sIC-Ne9BMwZZtBvR3FwKN_9a724704a550471a9df59999e9e1017f.webp)
Text-to-image generation with FLUX.2 [klein] 9B from Black Forest Labs and custom LoRA.

Run any LLM (Large Language Model) with fal, powered by OpenRouter.

Wan-2.1 is a text-to-video model that generates high-quality videos with high visual quality and motion diversity from text prompts

VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video

Generate character-consistent videos from reference images using PixVerse C1, with subject and background references.

Prompt-free object removal from an image and mask, erasing objects with their shadows and reflections and reconstructing the scene cleanly.

Extend Veo-Created Videos up to 30 seconds

Generate high-fidelity, design-ready images with precise typography, strong prompt alignment, and rich visual detail using Microsoft's flagship MAI Image 2.5 Pro.

FLUX Control LoRA Depth is a high-performance endpoint that uses a control image to transfer structure to the generated image, using a depth map.

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.

Enhances a given raster image using the 'creative upscale' tool, increasing image resolution, making the image sharper and cleaner.

MuseTalk is a real-time high quality audio-driven lip-syncing model. Use MuseTalk to animate a face with your own audio.

Precise camera position and angle control (rotation, zoom, vertical movement)

Generate 3D models from a single image using Tripo P1.

Audio-driven talking avatar generation powered by the SoulX-FlashTalk 14B model.

Generate videos from prompts and images using LTX Video-0.9.7 13B Distilled and custom LoRA

Luma Uni-1 Max Edit applies text-guided edits to a source image at maximum fidelity, holding the original structure while honoring reference images for precise, high-detail revisions.

Edit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.

US hosted version of ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.

Generate long videos from prompts and images using LTX Video-0.9.8 13B Distilled and custom LoRA

Image2SVG transforms raster images into clean vector graphics, preserving visual quality while enabling scalable, customizable SVG outputs with precise control over detail levels.

Image editing with HY-WU. Transfer outfits, swap faces, and blend textures instantly—no finetuning needed, just describe what you want and provide reference images.

Pixverse Effects

Removes mask-selected objects and their visual effects, seamlessly reconstructing the scene with contextually appropriate content.

Happy Horse 1.1 is Alibaba's #1-ranked video model. This text-to-video endpoint generates 1080p video with synchronized native audio and multilingual lip-sync from a text prompt alone.

Recraft V4.1 Pro Vector generates large-format, fully editable SVGs with the structural clarity professional illustrators expect. Built for poster art, complex brand assets, and detailed scene illustration, it scales without losing geometric integrity.

Wan-2.1 flf2v generates dynamic videos by intelligently bridging a given first frame to a desired end frame through smooth, coherent motion sequences.