Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Vision Models

AI Models with Vision: Multimodal LLMs for Image Understanding

Model rankings updated September 2026 based on real usage data.

Vision models are multimodal LLMs that analyze images, read documents, interpret charts, and answer questions about visual content alongside text. This collection ranks vision-capable models by their usage on OpenRouter over the past week. The current top models are GPT-5.6 Luna, GLM 5.3 Flash, and MiMo-V2.5. Access models from Anthropic, Google, OpenAI, and other providers through a single API, and compare context length, pricing, and capabilities.

Browse All ModelsCompare Models

Top Vision Models on OpenRouter

Favicon for openai

OpenAI: GPT-5.6 Luna

15.1T tokens
Academia (#4)
Finance (#4)
Health (#5)
Legal (#2)
Marketing (#2)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3 Flash

13.3T tokens
Academia (#2)
Finance (#2)
Health (#3)
Legal (#5)
Marketing (#4)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.31M context$0.075/M input tokens$0.25/M output tokens50% off
Favicon for xiaomi

Xiaomi: MiMo-V2.5

5.96T tokens
Academia (#11)
Finance (#15)
Health (#16)
Marketing (#19)
SEO (#24)

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for google

Google: Gemini 3.8 Flash

2.76T tokens
Academia (#12)
Finance (#10)
Health (#9)
Legal (#11)
Marketing (#15)

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

by google1.05M context$0.75/M input tokens$3.75/M output tokens50% off
Favicon for meta

Meta: Muse Spark 1.3 Contributor

2.05T tokens
Academia (#21)
Finance (#30)
Health (#44)
Legal (#38)
Marketing (#26)

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed. Prompts and outputs may be used to improve Meta’s products.

by meta1.05M context$0.10/M input tokens$0.20/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Sol

1.87T tokens
Academia (#14)
Finance (#26)
Health (#23)
Legal (#24)
Marketing (#7)

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks and long-horizon problem solving.

by openai1.05M context$2/M input tokens$10/M output tokens50% off
Favicon for moonshotai

MoonshotAI: Kimi K3

1.84T tokens
Academia (#34)
Finance (#21)
Health (#32)
Legal (#36)
Marketing (#42)

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

by moonshotai1.05M context$2.34/M input tokens$11.70/M output tokens
Favicon for anthropic

Claude Opus 5

1.82T tokens
Academia (#23)
Finance (#17)
Health (#22)
Legal (#42)
Marketing (#22)

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents.

The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

by anthropic1M context$5/M input tokens$25/M output tokens
Favicon for minimax

MiniMax: MiniMax M3

1.61T tokens
Academia (#25)
Finance (#11)
Health (#47)
Legal (#20)
Marketing (#33)

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

by minimax1.05M context$0.23/M input tokens$0.96/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4.1 Flash

1.55T tokens
Finance (#12)
Marketing (#44)
Programming (#22)
Roleplay (#47)
Science (#28)

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

by deepseek1.05M context$0.15/M input tokens$0.60/M output tokens

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Roleplay
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections