CritPt benchmark

349 AI models ranked on CritPt. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GPT-5.6 Sol (OpenAI) leads at 32%, from $2.00 per 1M input.

Data via arXiv:2509.26574

Models ranked

349

on this benchmark

Top score

32%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs CritPt

213 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on CritPt

highest score first · 349 models
Maker
#
1GPT-5.6 SolOpenAI32%$2.00$10.00
2GPT-5.5 ProOpenAI31%$30.00$180.00
3GPT-5.6 TerraOpenAI30%$2.00$12.00
4GPT-5.4 ProOpenAI30%$30.00$180.00
5Claude Opus 5Anthropic29%$5.00$25.00
6Claude Fable 5Anthropic29%$10.00$50.00
7GPT-5.5OpenAI27%$5.00$30.00
8Gemini 3 Deep ThinkGoogle26%n/an/a
9GPT-5.4OpenAI23%$2.50$15.00
10Kimi K3Moonshot23%$2.85$14.25
11GLM 5.2Zhipu21%$0.61$1.98
12Claude Opus 4.8Anthropic21%$5.00$25.00
13GPT-5.6 LunaOpenAI21%$0.20$1.20
14Qwen3.8 MaxAlibaba20%$1.65$4.95
15Qwen3.8 2.4T A95BAlibaba20%$2.00$6.00
16GLM 5.3Zhipu19%$1.40$4.40
17DeepSeek V4 Pro 0423DeepSeek18%$0.43$0.87
18Gemini 3.1 Pro PreviewGoogle18%$2.00$12.00
19Muse Spark 1.2Other18%$1.25$4.25
20Grok 4.6xAI17%$2.00$6.00
21Claude Sonnet 5Anthropic17%$2.00$10.00
22GPT-5.3-CodexOpenAI17%$1.75$14.00
23DeepSeek V4 Flash 0423DeepSeek17%$0.09$0.17
24Agnes 2.5 Pro BetaSapiens-ai16%n/an/a
25GLM 5.3 FlashZhipu15%$0.07$0.25
26Grok 4.5xAI15%$2.00$6.00
27Muse Spark 1.1Other15%$1.25$4.25
28Gemini 3.7 FlashGoogle14%$0.38$1.88
29Qwen3.7 MaxAlibaba13%$1.25$3.75
30Gemini 3.5 FlashGoogle13%$1.50$9.00
31DeepSeek V4 Pro (max)DeepSeek13%n/an/a
32Claude Opus 4.6Anthropic13%reasoning: adaptive$5.00$25.00
33Claude Opus 4.7Anthropic12%$5.00$25.00
34GPT-5.2OpenAI12%$1.75$14.00
35Muse SparkMeta11%n/an/a
36Qwen3.8-Flash-NextAlibaba11%n/an/a
37Agnes 2.5 Pro AlphaSapiens-ai11%n/an/a
38DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek11%n/an/a
39Gemini 3.6 FlashGoogle11%$0.75$3.75
40Kimi K2.7 CodeMoonshot10%$0.66$3.40
41GPT-5.4 MiniOpenAI10%$0.75$4.50
42GPT-5.4 NanoOpenAI9%$0.20$1.25
43Qwen3.7 PlusAlibaba9%$0.32$1.28
44Gemini 3 ProGoogle9%$2.00$12.00
45Grok Build 0.1 0616Spacexai9%n/an/a
46A.X-K2Sk-telecom9%n/an/a
47GPT-5.2-CodexOpenAI9%$1.75$14.00
48Nex-N2-ProOther9%$0.25$1.00
49Gemini 3 FlashGoogle9%reasoning: true$0.63$3.75
50Inkling SmallOther8%$0.45$1.20
150 of 349
Page 1 of 7
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the CritPt benchmark?

CritPt (Complex Research using Integrated Thinking, Physics Test) is a frontier physics benchmark of 71 unpublished, research-level challenges across 12 subdisciplines, created by more than 50 active physicists. Scored on accuracy in guess-resistant formats (numeric arrays, symbolic expressions, code), where the best current models solve only single-digit percentages unaided.

Which AI model scores highest on CritPt?

GPT-5.6 Sol (OpenAI) leads with 32%, from $2.00 per 1M input tokens.

How many models are ranked on CritPt?

349 models carry a CritPt score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI