
Upscale videos to 1080p, 2K, or 4K via API. FLUX 3 powered super-resolution with a precise mode and a creative detail-enhancement mode.

SAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.

Run any audio capable LLM with fal. Process audio files — transcription, analysis, understanding, understand— using Gemini (Google) models. Supports wav, mp3, aiff, aac, ogg, flac, m4a. Powered by OpenRouter.

Use Gemini TTS Models to convert your prompts to real audio.

Turn any flat ad image into fully editable layers —background, product and logo cutouts, live text with typography, and vector shapes. Commercial-safe, structured JSON output

Recraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy — delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.

MiniMax Hailuo-2.3 Image To Video API (Standard, 768p): Advanced image-to-video generation model with 768p resolution

MAI-Image-2.5 is Microsoft's photorealistic image generation and editing model that turns text prompts or uploaded images into high-quality, design-ready visuals with fine-grained, pixel-level control.

Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.

Generate video clips from your images using Kling 1.0

MiniMax Hailuo-02 Text To Video API (Standard, 768p): Advanced video generation model with 768p resolution

Kling 2.1 Master: The premium endpoint for Kling 2.1, designed for top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.

Rig humanoid 3D models from GLB URLs with Meshy, returning rigged GLB/FBX files plus basic animations.

Generate high fidelity, studio quality videos of your avatar speaking or singing using the Aurora from Creatify team!

Generate high-quality images from text with Krea 2 Medium, supporting aspect ratio, creativity controls, seeds, and optional style references.

FFMPEG Utilities to Scale Videos

Happy Horse 1.1 is Alibaba's #1-ranked video model. This image-to-video endpoint animates a still image into 1080p video with synchronized native audio and multilingual lip-sync

Image editing endpoint for Hunyuan Image 3.0 Instruct.

Wan 2.2's 5B model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

Change the voices in your audios with voices in ElevenLabs!

Professional creative image upscaling powered by Topaz Labs. Bloom 2 reinvents detail with adjustable creativity and color preservation. Best for AI-generated images that need striking enhancement.

Rodin by Hyper3D generates realistic and production ready 3D models from text or images.

LTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

Generate complete seamlessly tiling PBR materials including normal, roughness, basecolor, height and metalness maps up to 8K

OmniHuman generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.

Generate consistent character appearances across multiple images. Maintain facial features, proportions, and distinctive traits for cohesive storytelling and branding

Run SDXL at the speed of light