Skip to main content

Overview

Models for understanding and analyzing images, including captioning, visual question answering, and object detection.

Top Models

OpenRouter [Vision] API

Run any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. S
Example output from OpenRouter [Vision]

NSFW Filter API

Predict the probability of an image being NSFW.
Example output from NSFW Filter

Florence-2 Large API

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
Example output from Florence-2 Large
Explore all vision models on fal.ai/models.

Quick Start

Get started with OpenRouter [Vision]:

Pricing

For detailed pricing information, see the fal.ai pricing page or individual model pages.