GDPval benchmark
162 AI models ranked on GDPval. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Claude Opus 5 (Anthropic) leads at 1844.7, from $5.00 per 1M input.
Data via arXiv:2510.04374 →Models ranked
162
on this benchmark
Top score
1844.7
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Models ranked on GDPval
highest score first · 162 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 1844.7 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 2 | Grok 4.6 | 1747.5 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 3 | Claude Fable 5 | 1737.7 | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 4 | Qwen3.8 Max | 1735.2 | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 5 | GPT-5.6 Sol | 1723.0 | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 6 | Qwen3.8 2.4T A95B | 1720.4 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 7 | Kimi K3 | 1681.1 | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 8 | Muse Spark 1.2 | Other | 1628 | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 9 | Claude Sonnet 5 | 1595.5 | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | DeepSeek V4 Pro 0423 | 1590.4 | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 11 | Claude Opus 4.8 | 1583.7 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 12 | GPT-5.6 Luna | 1578.3 | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 13 | GPT-5.6 Terra | 1576.3 | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | DeepSeek V4 Flash 0423 | 1558.9 | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 15 | Qwen3.8 27B | 1545.9 | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 16 | Gemini 3.7 Flash | 1531.5 | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 17 | Grok 4.5 | 1524.4 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 18 | GLM 5.2 | 1505.2 | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 19 | GPT-5.5 | 1490.2 | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | Claude Opus 4.7 | 1489.5 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 21 | Gemini 3.6 Flash | 1422.2 | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 22 | GPT-5.4 | 1394.0 | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| 23 | MiniMax M3 | 1386.9 | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 24 | Claude Sonnet 4.6 | 1374.3reasoning: adaptive | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| 25 | Muse Spark 1.1 | Other | 1374.2 | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 26 | Gemini 3.5 Flash | 1343.4 | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| 27 | DeepSeek V4 Pro (max) | 1306.1 | n/a | n/a | ||||||||||||||||||||||||||||||
| 28 | Solar Pro 4 | Other | 1275.8 | $0.03 | $0.12 | |||||||||||||||||||||||||||||
| 29 | Motif 3 | Motif-technologies | 1274.5 | n/a | n/a | |||||||||||||||||||||||||||||
| 30 | Qwen3.7 Max | 1271 | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 31 | Inkling Small | Other | 1268.2 | $0.45 | $1.20 | |||||||||||||||||||||||||||||
| 32 | MiMo-V2.5-Pro | 1265.8 | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| 33 | JT-4.1 Flash 236B A21B | China-mobile | 1263.4 | n/a | n/a | |||||||||||||||||||||||||||||
| 34 | GLM 5.1 | 1257.8 | $1.05 | $3.50 | ||||||||||||||||||||||||||||||
| 35 | Motif 3 (Beta) | Motif-technologies | 1256.7 | n/a | n/a | |||||||||||||||||||||||||||||
| 36 | Nex-N2-Pro | Other | 1248.9 | $0.25 | $1.00 | |||||||||||||||||||||||||||||
| 37 | Inkling | Other | 1238.7 | $0.95 | $4.05 | |||||||||||||||||||||||||||||
| 38 | Grok Build 0.1 0616 | Spacexai | 1215.1 | n/a | n/a | |||||||||||||||||||||||||||||
| 39 | Hy3 | Other | 1214.4 | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| 40 | G9v3-39A5B | Ai9stars | 1196.6 | n/a | n/a | |||||||||||||||||||||||||||||
| 41 | Kimi K2.6 | 1190.4 | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| 42 | DeepSeek V4 Flash (max) | 1190.0 | n/a | n/a | ||||||||||||||||||||||||||||||
| 43 | Kimi K2.7 Code | 1189.9 | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 44 | Agnes 2.5 Pro Alpha | Sapiens-ai | 1175.6 | n/a | n/a | |||||||||||||||||||||||||||||
| 45 | GPT-5.4 Mini | 1171.8 | $0.75 | $4.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 46 | GLM 4.7 | 1169.1 | $0.40 | $1.75 | ||||||||||||||||||||||||||||||
| 47 | Nvidia Nemotron 3 Ultra 550B A55b | 1163.0 | $0.50 | $2.20 | ||||||||||||||||||||||||||||||
| 48 | MiniMax M2.7 | 1159.8 | $0.25 | $0.55 | ||||||||||||||||||||||||||||||
| 49 | MiMo-V2.5 | 1150.4 | n/a | n/a | ||||||||||||||||||||||||||||||
| 50 | Muse Spark | 1146.3 | n/a | n/a | ||||||||||||||||||||||||||||||
Frequently asked questions
What is the GDPval benchmark?
GDPval (OpenAI, 2025) evaluates AI on 1,320 real-world knowledge-work tasks drawn from 44 occupations across the nine largest US GDP sectors, built from the work of experienced professionals. Expert graders blind-compare model deliverables against professional work, producing win or tie rates.
Which AI model scores highest on GDPval?
Claude Opus 5 (Anthropic) leads with 1844.7, from $5.00 per 1M input tokens.
How many models are ranked on GDPval?
162 models carry a GDPval score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI