CritPt benchmark
349 AI models ranked on CritPt. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
GPT-5.6 Sol (OpenAI) leads at 32%, from $2.00 per 1M input.
Data via arXiv:2509.26574 →Models ranked
349
on this benchmark
Top score
32%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs CritPt
213 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on CritPt
highest score first · 349 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 32% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 2 | GPT-5.5 Pro | 31% | $30.00 | $180.00 | ||||||||||||||||||||||||||||||
| 3 | GPT-5.6 Terra | 30% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 4 | GPT-5.4 Pro | 30% | $30.00 | $180.00 | ||||||||||||||||||||||||||||||
| 5 | Claude Opus 5 | 29% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 6 | Claude Fable 5 | 29% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 7 | GPT-5.5 | 27% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 8 | Gemini 3 Deep Think | 26% | n/a | n/a | ||||||||||||||||||||||||||||||
| 9 | GPT-5.4 | 23% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | Kimi K3 | 23% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 11 | GLM 5.2 | 21% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 12 | Claude Opus 4.8 | 21% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 13 | GPT-5.6 Luna | 21% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | Qwen3.8 Max | 20% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 15 | Qwen3.8 2.4T A95B | 20% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 16 | GLM 5.3 | 19% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 17 | DeepSeek V4 Pro 0423 | 18% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 18 | Gemini 3.1 Pro Preview | 18% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 19 | Muse Spark 1.2 | Other | 18% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 20 | Grok 4.6 | 17% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 21 | Claude Sonnet 5 | 17% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | GPT-5.3-Codex | 17% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 23 | DeepSeek V4 Flash 0423 | 17% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | Agnes 2.5 Pro Beta | Sapiens-ai | 16% | n/a | n/a | |||||||||||||||||||||||||||||
| 25 | GLM 5.3 Flash | 15% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 26 | Grok 4.5 | 15% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 27 | Muse Spark 1.1 | Other | 15% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 28 | Gemini 3.7 Flash | 14% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 29 | Qwen3.7 Max | 13% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 30 | Gemini 3.5 Flash | 13% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | DeepSeek V4 Pro (max) | 13% | n/a | n/a | ||||||||||||||||||||||||||||||
| 32 | Claude Opus 4.6 | 13%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 33 | Claude Opus 4.7 | 12% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 34 | GPT-5.2 | 12% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 35 | Muse Spark | 11% | n/a | n/a | ||||||||||||||||||||||||||||||
| 36 | Qwen3.8-Flash-Next | 11% | n/a | n/a | ||||||||||||||||||||||||||||||
| 37 | Agnes 2.5 Pro Alpha | Sapiens-ai | 11% | n/a | n/a | |||||||||||||||||||||||||||||
| 38 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 11% | n/a | n/a | ||||||||||||||||||||||||||||||
| 39 | Gemini 3.6 Flash | 11% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 40 | Kimi K2.7 Code | 10% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 41 | GPT-5.4 Mini | 10% | $0.75 | $4.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 42 | GPT-5.4 Nano | 9% | $0.20 | $1.25 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 43 | Qwen3.7 Plus | 9% | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
| 44 | Gemini 3 Pro | 9% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Grok Build 0.1 0616 | Spacexai | 9% | n/a | n/a | |||||||||||||||||||||||||||||
| 46 | A.X-K2 | Sk-telecom | 9% | n/a | n/a | |||||||||||||||||||||||||||||
| 47 | GPT-5.2-Codex | 9% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 48 | Nex-N2-Pro | Other | 9% | $0.25 | $1.00 | |||||||||||||||||||||||||||||
| 49 | Gemini 3 Flash | 9%reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 50 | Inkling Small | Other | 8% | $0.45 | $1.20 | |||||||||||||||||||||||||||||
Frequently asked questions
What is the CritPt benchmark?
CritPt (Complex Research using Integrated Thinking, Physics Test) is a frontier physics benchmark of 71 unpublished, research-level challenges across 12 subdisciplines, created by more than 50 active physicists. Scored on accuracy in guess-resistant formats (numeric arrays, symbolic expressions, code), where the best current models solve only single-digit percentages unaided.
Which AI model scores highest on CritPt?
GPT-5.6 Sol (OpenAI) leads with 32%, from $2.00 per 1M input tokens.
How many models are ranked on CritPt?
349 models carry a CritPt score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI