openrouter/router/vision
Run any Vision Language Model with fal. Analyze and understand images using Claude (Anthropic), GPT-5 / GPT-4o (OpenAI), Gemini (Google), Grok (xAI), Llama (Meta), Qwen, Pixtral (Mistral), and more. Send one or multiple images for captioning, analysis, OCR, or visual Q&A. Powered by OpenRouter.
Inference
Commercial use
Streaming
Partner
Input
Hint: Drag and drop files from your computer, images from web pages, paste from clipboard (Ctrl/Cmd+V), or provide a URL.
1 image added
Type # to reference inputs.
Additional Settings
Customize your input with more control.
Result
Idle
What would you like to do next?
You will be charged based on the number of input and output tokens.
