GPQA Diamond benchmark
441 AI models ranked on GPQA Diamond. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Models ranked
441
on this benchmark
Top score
95%
current leader
Type
Single eval
benchmark
Updated
30 Aug 2026
last source fetch
Cost vs GPQA Diamond
263 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on GPQA Diamond
highest score first · 441 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Grok 4.6 | 95% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 2 | Gemini 3.7 Flash | 95% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 3 | Gemini 3.1 Pro Preview | 94% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 4 | GPT-5.6 Sol | 94% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 5 | Kimi K3 | 94% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 6 | Qwen3.8 2.4T A95B | 94% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 7 | GPT-5.5 | 94% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 8 | Claude Opus 5 | 93% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 9 | Grok 4.5 | 93% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 10 | MiniMax M3 | 93% | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 11 | DeepSeek V4 Pro 0423 | 93% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 12 | Gemini 3.6 Flash | 93% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 13 | Qwen3.8 Max | 93% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 14 | Claude Fable 5 | 93% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 15 | GPT-5.6 Terra | 93% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 16 | Qwen3.7 Max | 92% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 17 | Qwen3.8-Flash-Next | 92% | n/a | n/a | ||||||||||||||||||||||||||||||
| 18 | Gemini 3.5 Flash | 92% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 19 | Claude Opus 4.8 | 92% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 20 | GPT-5.4 | 92% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 21 | GLM 5.3 | 92% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 22 | GPT-5.3-Codex | 92% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 23 | Claude Opus 4.7 | 91% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 91% | n/a | n/a | ||||||||||||||||||||||||||||||
| 25 | GLM 5.3 Flash | 91% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 26 | Kimi K2.6 | 91% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 27 | Claude Sonnet 5 | 91% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 28 | GPT-5.6 Luna | 91% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 29 | Grok 4.20 | 91% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| 30 | DeepSeek V4 Flash 0423 | 91% | $0.08 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | Gemini 3 Pro | 91% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 32 | Qwen3.8 27B | 91% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 33 | Agnes 2.5 Pro Beta | Sapiens-ai | 91% | n/a | n/a | |||||||||||||||||||||||||||||
| 34 | Muse Spark 1.2 | Other | 90% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 35 | GPT-5.2 | 90% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 36 | Grok 4.3 | 90% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | Qwen3.7 Plus | 90% | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
| 38 | GPT-5.2-Codex | 90% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 39 | Muse Spark 1.1 | Other | 90% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 40 | Gemini 3 Flash | 90%reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 41 | Hy3 | Other | 90% | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 42 | Kimi K2.7 Code | 90% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 43 | Claude Opus 4.6 | 90%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 44 | GLM 5.2 | 89% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Inkling Small | Other | 89% | $0.45 | $1.20 | |||||||||||||||||||||||||||||
| 46 | Grok Build 0.1 0616 | Spacexai | 89% | n/a | n/a | |||||||||||||||||||||||||||||
| 47 | DeepSeek V4 Flash (max) | 89% | n/a | n/a | ||||||||||||||||||||||||||||||
| 48 | Qwen3.5 397B A17B | 89% | $0.39 | $2.34 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 49 | Nex-N2-Pro | Other | 89% | $0.25 | $1.00 | |||||||||||||||||||||||||||||
| 50 | Solar Pro 4 | Other | 89% | $0.03 | $0.12 | |||||||||||||||||||||||||||||
Where a model has no directly-measured score, a lab-claimed or estimated value is shown.
Frequently asked questions
What is the GPQA Diamond benchmark?
A 198-question multiple-choice subset of GPQA covering graduate-level biology, physics and chemistry, written to be 'Google-proof' so answers can't be found by web search. Scored on accuracy: domain PhDs reach about 65%, while skilled non-experts manage only about 34% even with web access.
Which AI model scores highest on GPQA Diamond?
Grok 4.6 (xAI) leads with 95%, from $2.00 per 1M input tokens.
How many models are ranked on GPQA Diamond?
441 models carry a GPQA Diamond score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI