GDPval benchmark

162 AI models ranked on GDPval. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Claude Opus 5 (Anthropic) leads at 1844.7, from $5.00 per 1M input.

Data via arXiv:2510.04374

Models ranked

162

on this benchmark

Top score

1844.7

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Models ranked on GDPval

highest score first · 162 models
Maker
#
1Claude Opus 5Anthropic1844.7$5.00$25.00
2Grok 4.6xAI1747.5$2.00$6.00
3Claude Fable 5Anthropic1737.7$10.00$50.00
4Qwen3.8 MaxAlibaba1735.2$1.65$4.95
5GPT-5.6 SolOpenAI1723.0$2.00$10.00
6Qwen3.8 2.4T A95BAlibaba1720.4$2.00$6.00
7Kimi K3Moonshot1681.1$2.85$14.25
8Muse Spark 1.2Other1628$1.25$4.25
9Claude Sonnet 5Anthropic1595.5$2.00$10.00
10DeepSeek V4 Pro 0423DeepSeek1590.4$0.43$0.87
11Claude Opus 4.8Anthropic1583.7$5.00$25.00
12GPT-5.6 LunaOpenAI1578.3$0.20$1.20
13GPT-5.6 TerraOpenAI1576.3$2.00$12.00
14DeepSeek V4 Flash 0423DeepSeek1558.9$0.09$0.17
15Qwen3.8 27BAlibaba1545.9$0.40$2.55
16Gemini 3.7 FlashGoogle1531.5$0.38$1.88
17Grok 4.5xAI1524.4$2.00$6.00
18GLM 5.2Zhipu1505.2$0.61$1.98
19GPT-5.5OpenAI1490.2$5.00$30.00
20Claude Opus 4.7Anthropic1489.5$5.00$25.00
21Gemini 3.6 FlashGoogle1422.2$0.75$3.75
22GPT-5.4OpenAI1394.0$2.50$15.00
23MiniMax M3MiniMax1386.9$0.23$0.96
24Claude Sonnet 4.6Anthropic1374.3reasoning: adaptive$3.00$15.00
25Muse Spark 1.1Other1374.2$1.25$4.25
26Gemini 3.5 FlashGoogle1343.4$1.50$9.00
27DeepSeek V4 Pro (max)DeepSeek1306.1n/an/a
28Solar Pro 4Other1275.8$0.03$0.12
29Motif 3Motif-technologies1274.5n/an/a
30Qwen3.7 MaxAlibaba1271$1.25$3.75
31Inkling SmallOther1268.2$0.45$1.20
32MiMo-V2.5-ProXiaomi1265.8$0.43$0.87
33JT-4.1 Flash 236B A21BChina-mobile1263.4n/an/a
34GLM 5.1Zhipu1257.8$1.05$3.50
35Motif 3 (Beta)Motif-technologies1256.7n/an/a
36Nex-N2-ProOther1248.9$0.25$1.00
37InklingOther1238.7$0.95$4.05
38Grok Build 0.1 0616Spacexai1215.1n/an/a
39Hy3Other1214.4$0.13$0.53
40G9v3-39A5BAi9stars1196.6n/an/a
41Kimi K2.6Moonshot1190.4$0.65$3.40
42DeepSeek V4 Flash (max)DeepSeek1190.0n/an/a
43Kimi K2.7 CodeMoonshot1189.9$0.66$3.40
44Agnes 2.5 Pro AlphaSapiens-ai1175.6n/an/a
45GPT-5.4 MiniOpenAI1171.8$0.75$4.50
46GLM 4.7Zhipu1169.1$0.40$1.75
47Nvidia Nemotron 3 Ultra 550B A55bNVIDIA1163.0$0.50$2.20
48MiniMax M2.7MiniMax1159.8$0.25$0.55
49MiMo-V2.5Xiaomi1150.4n/an/a
50Muse SparkMeta1146.3n/an/a
150 of 162
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the GDPval benchmark?

GDPval (OpenAI, 2025) evaluates AI on 1,320 real-world knowledge-work tasks drawn from 44 occupations across the nine largest US GDP sectors, built from the work of experienced professionals. Expert graders blind-compare model deliverables against professional work, producing win or tie rates.

Which AI model scores highest on GDPval?

Claude Opus 5 (Anthropic) leads with 1844.7, from $5.00 per 1M input tokens.

How many models are ranked on GDPval?

162 models carry a GDPval score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI