MMMU-Pro benchmark
151 AI models ranked on MMMU-Pro. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Gemini 3.7 Flash (Google) leads at 85%, from $0.38 per 1M input.
Data via arXiv:2409.02813 →Models ranked
151
on this benchmark
Top score
85%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs MMMU-Pro
106 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on MMMU-Pro
highest score first · 151 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | 85% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 2 | Claude Opus 5 | 85% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 3 | Gemini 3.5 Flash | 84% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 4 | GPT-5.6 Sol | 83% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 5 | Gemini 3.6 Flash | 83% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 6 | Gemini 3.1 Pro Preview | 82% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 7 | Qwen3.8 Max | 82% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 8 | GPT-5.5 | 81%medium | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 9 | GPT-5.6 Terra | 81% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | Kimi K3 | 81% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 11 | Muse Spark | 81% | n/a | n/a | ||||||||||||||||||||||||||||||
| 12 | Qwen3.7 Plus | 80% | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
| 13 | Grok 4.5 | 80% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 14 | Gemini 3 Pro | 80% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 15 | Gemini 3 Flash | 80%reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 16 | Qwen3.8-Flash-Next | 80% | n/a | n/a | ||||||||||||||||||||||||||||||
| 17 | Kimi K2.6 | 79% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| 18 | Gemini 3.5 Flash Lite | 79% | $0.30 | $2.50 | ||||||||||||||||||||||||||||||
| 19 | Claude Opus 4.7 | 79% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | MiniMax M3 | 79% | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 21 | GPT-5.6 Luna | 79% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | GPT-5.3-Codex | 78% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 23 | GPT-5.4 | 78% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | Grok 4.3 | 78% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 25 | Qwen3.6 Plus | 78% | $0.33 | $1.95 | ||||||||||||||||||||||||||||||
| 26 | Claude Sonnet 5 | 77% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 27 | Qwen3.5 397B A17B | 77% | $0.39 | $2.34 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 28 | Grok Build 0.1 0616 | Spacexai | 77% | n/a | n/a | |||||||||||||||||||||||||||||
| 29 | Qwen3.8 27B | 76% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 30 | GPT-5.2-Codex | 76% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 31 | Gemini 3.1 Flash Lite Preview | 76% | $0.25 | $1.50 | ||||||||||||||||||||||||||||||
| 32 | GPT-5.1 | 75% | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 33 | Claude Opus 4.6 | 75%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 34 | Agnes 2.5 Pro Beta | Sapiens-ai | 75% | n/a | n/a | |||||||||||||||||||||||||||||
| 35 | MiMo-V2.5 | 75% | n/a | n/a | ||||||||||||||||||||||||||||||
| 36 | Kimi K2.5 | 75% | $0.45 | $2.25 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | Step 3.7 Flash | Other | 75% | $0.20 | $1.15 | |||||||||||||||||||||||||||||
| 38 | Qwen3.6 35B A3B | 75% | $0.10 | $0.45 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 39 | Qwen3.5-27B | 75% | $0.20 | $1.56 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 40 | Qwen3.5-122B-A10B | 75% | $0.25 | $1.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 41 | Gemini 2.5 Pro | 75% | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| 42 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 75% | n/a | n/a | ||||||||||||||||||||||||||||||
| 43 | Qwen3.6 27B | 75% | $0.15 | $0.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 44 | GPT-5.2 | 75%medium | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Grok 4.20 | 75% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| 46 | GPT-5 | 74%medium | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 47 | Muse Glimmer | 74% | n/a | n/a | ||||||||||||||||||||||||||||||
| 48 | Claude Opus 4.5 | 74%reasoning: true | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 49 | Inkling Small | Other | 74% | $0.45 | $1.20 | |||||||||||||||||||||||||||||
| 50 | MiMo-V2-Omni-0327 | 74% | n/a | n/a | ||||||||||||||||||||||||||||||
Frequently asked questions
What is the MMMU-Pro benchmark?
A more robust version of the MMMU multimodal benchmark testing reasoning over combined images and text across many disciplines. It filters out text-only-answerable questions, expands the answer options, and adds a vision-only mode where the question is embedded in the image. Scored as multiple-choice accuracy.
Which AI model scores highest on MMMU-Pro?
Gemini 3.7 Flash (Google) leads with 85%, from $0.38 per 1M input tokens.
How many models are ranked on MMMU-Pro?
151 models carry a MMMU-Pro score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI