MMMU-Pro benchmark

151 AI models ranked on MMMU-Pro. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Gemini 3.7 Flash (Google) leads at 85%, from $0.38 per 1M input.

Data via arXiv:2409.02813

Models ranked

151

on this benchmark

Top score

85%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs MMMU-Pro

106 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on MMMU-Pro

highest score first · 151 models
Maker
#
1Gemini 3.7 FlashGoogle85%$0.38$1.88
2Claude Opus 5Anthropic85%$5.00$25.00
3Gemini 3.5 FlashGoogle84%$1.50$9.00
4GPT-5.6 SolOpenAI83%$2.00$10.00
5Gemini 3.6 FlashGoogle83%$0.75$3.75
6Gemini 3.1 Pro PreviewGoogle82%$2.00$12.00
7Qwen3.8 MaxAlibaba82%$1.65$4.95
8GPT-5.5OpenAI81%medium$5.00$30.00
9GPT-5.6 TerraOpenAI81%$2.00$12.00
10Kimi K3Moonshot81%$2.85$14.25
11Muse SparkMeta81%n/an/a
12Qwen3.7 PlusAlibaba80%$0.32$1.28
13Grok 4.5xAI80%$2.00$6.00
14Gemini 3 ProGoogle80%$2.00$12.00
15Gemini 3 FlashGoogle80%reasoning: true$0.63$3.75
16Qwen3.8-Flash-NextAlibaba80%n/an/a
17Kimi K2.6Moonshot79%$0.65$3.40
18Gemini 3.5 Flash LiteGoogle79%$0.30$2.50
19Claude Opus 4.7Anthropic79%$5.00$25.00
20MiniMax M3MiniMax79%$0.23$0.96
21GPT-5.6 LunaOpenAI79%$0.20$1.20
22GPT-5.3-CodexOpenAI78%$1.75$14.00
23GPT-5.4OpenAI78%$2.50$15.00
24Grok 4.3xAI78%$1.25$2.50
25Qwen3.6 PlusAlibaba78%$0.33$1.95
26Claude Sonnet 5Anthropic77%$2.00$10.00
27Qwen3.5 397B A17BAlibaba77%$0.39$2.34
28Grok Build 0.1 0616Spacexai77%n/an/a
29Qwen3.8 27BAlibaba76%$0.40$2.55
30GPT-5.2-CodexOpenAI76%$1.75$14.00
31Gemini 3.1 Flash Lite PreviewGoogle76%$0.25$1.50
32GPT-5.1OpenAI75%$1.25$10.00
33Claude Opus 4.6Anthropic75%reasoning: adaptive$5.00$25.00
34Agnes 2.5 Pro BetaSapiens-ai75%n/an/a
35MiMo-V2.5Xiaomi75%n/an/a
36Kimi K2.5Moonshot75%$0.45$2.25
37Step 3.7 FlashOther75%$0.20$1.15
38Qwen3.6 35B A3BAlibaba75%$0.10$0.45
39Qwen3.5-27BAlibaba75%$0.20$1.56
40Qwen3.5-122B-A10BAlibaba75%$0.25$1.75
41Gemini 2.5 ProGoogle75%$1.25$10.00
42DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek75%n/an/a
43Qwen3.6 27BAlibaba75%$0.15$0.50
44GPT-5.2OpenAI75%medium$1.75$14.00
45Grok 4.20xAI75%$1.25$2.50
46GPT-5OpenAI74%medium$1.25$10.00
47Muse GlimmerMeta74%n/an/a
48Claude Opus 4.5Anthropic74%reasoning: true$5.00$25.00
49Inkling SmallOther74%$0.45$1.20
50MiMo-V2-Omni-0327Xiaomi74%n/an/a
150 of 151
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the MMMU-Pro benchmark?

A more robust version of the MMMU multimodal benchmark testing reasoning over combined images and text across many disciplines. It filters out text-only-answerable questions, expands the answer options, and adds a vision-only mode where the question is embedded in the image. Scored as multiple-choice accuracy.

Which AI model scores highest on MMMU-Pro?

Gemini 3.7 Flash (Google) leads with 85%, from $0.38 per 1M input tokens.

How many models are ranked on MMMU-Pro?

151 models carry a MMMU-Pro score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI