GPQA Diamond benchmark

441 AI models ranked on GPQA Diamond. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Grok 4.6 (xAI) leads at 95%, from $2.00 per 1M input.

Data via arXiv:2311.12022

Models ranked

441

on this benchmark

Top score

95%

current leader

Type

Single eval

benchmark

Updated

30 Aug 2026

last source fetch

Cost vs GPQA Diamond

263 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on GPQA Diamond

highest score first · 441 models
Maker
#
1Grok 4.6xAI95%$2.00$6.00
2Gemini 3.7 FlashGoogle95%$0.75$3.75
3Gemini 3.1 Pro PreviewGoogle94%$2.00$12.00
4GPT-5.6 SolOpenAI94%$2.00$10.00
5Kimi K3Moonshot94%$2.85$14.25
6Qwen3.8 2.4T A95BAlibaba94%$2.00$6.00
7GPT-5.5OpenAI94%$5.00$30.00
8Claude Opus 5Anthropic93%$5.00$25.00
9Grok 4.5xAI93%$2.00$6.00
10MiniMax M3MiniMax93%$0.23$0.96
11DeepSeek V4 Pro 0423DeepSeek93%$0.43$0.87
12Gemini 3.6 FlashGoogle93%$0.75$3.75
13Qwen3.8 MaxAlibaba93%$1.65$4.95
14Claude Fable 5Anthropic93%$10.00$50.00
15GPT-5.6 TerraOpenAI93%$2.00$12.00
16Qwen3.7 MaxAlibaba92%$1.25$3.75
17Qwen3.8-Flash-NextAlibaba92%n/an/a
18Gemini 3.5 FlashGoogle92%$1.50$9.00
19Claude Opus 4.8Anthropic92%$5.00$25.00
20GPT-5.4OpenAI92%$2.50$15.00
21GLM 5.3Zhipu92%$1.40$4.40
22GPT-5.3-CodexOpenAI92%$1.75$14.00
23Claude Opus 4.7Anthropic91%$5.00$25.00
24DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek91%n/an/a
25GLM 5.3 FlashZhipu91%$0.07$0.25
26Kimi K2.6Moonshot91%$0.65$3.40
27Claude Sonnet 5Anthropic91%$2.00$10.00
28GPT-5.6 LunaOpenAI91%$0.20$1.20
29Grok 4.20xAI91%$1.25$2.50
30DeepSeek V4 Flash 0423DeepSeek91%$0.08$0.17
31Gemini 3 ProGoogle91%$2.00$12.00
32Qwen3.8 27BAlibaba91%$0.40$2.55
33Agnes 2.5 Pro BetaSapiens-ai91%n/an/a
34Muse Spark 1.2Other90%$1.25$4.25
35GPT-5.2OpenAI90%$1.75$14.00
36Grok 4.3xAI90%$1.25$2.50
37Qwen3.7 PlusAlibaba90%$0.32$1.28
38GPT-5.2-CodexOpenAI90%$1.75$14.00
39Muse Spark 1.1Other90%$1.25$4.25
40Gemini 3 FlashGoogle90%reasoning: true$0.63$3.75
41Hy3Other90%$0.13$0.53
42Kimi K2.7 CodeMoonshot90%$0.66$3.40
43Claude Opus 4.6Anthropic90%reasoning: adaptive$5.00$25.00
44GLM 5.2Zhipu89%$0.61$1.98
45Inkling SmallOther89%$0.45$1.20
46Grok Build 0.1 0616Spacexai89%n/an/a
47DeepSeek V4 Flash (max)DeepSeek89%n/an/a
48Qwen3.5 397B A17BAlibaba89%$0.39$2.34
49Nex-N2-ProOther89%$0.25$1.00
50Solar Pro 4Other89%$0.03$0.12
150 of 441
Page 1 of 9
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Where a model has no directly-measured score, a lab-claimed or estimated value is shown.

Frequently asked questions

What is the GPQA Diamond benchmark?

A 198-question multiple-choice subset of GPQA covering graduate-level biology, physics and chemistry, written to be 'Google-proof' so answers can't be found by web search. Scored on accuracy: domain PhDs reach about 65%, while skilled non-experts manage only about 34% even with web access.

Which AI model scores highest on GPQA Diamond?

Grok 4.6 (xAI) leads with 95%, from $2.00 per 1M input tokens.

How many models are ranked on GPQA Diamond?

441 models carry a GPQA Diamond score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 30 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI