Humanity's Last Exam benchmark
430 AI models ranked on Humanity's Last Exam. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Claude Fable 5 (Anthropic) leads at 55%, from $10.00 per 1M input.
Data via arXiv:2501.14249 →Models ranked
430
on this benchmark
Top score
55%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs Humanity's Last Exam
256 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on Humanity's Last Exam
highest score first · 430 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 55% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 2 | Claude Opus 5 | 55% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 3 | GPT-5.6 Sol | 49% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 4 | Claude Opus 4.8 | 49% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 5 | Gemini 3.7 Flash | 48% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 6 | Gemini 3.1 Pro Preview | 47% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 7 | Kimi K3 | 47% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 8 | Muse Spark 1.1 | Other | 46% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 9 | GPT-5.5 | 46% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | Muse Spark 1.2 | Other | 45% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 11 | GPT-5.4 | 44% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 12 | Qwen3.8 Max | 43% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 13 | GPT-5.6 Terra | 43% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | Grok 4.6 | 43% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 15 | Gemini 3.5 Flash | 43% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 16 | Grok 4.5 | 43% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 17 | GPT-5.3-Codex | 42% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 18 | Qwen3.8 2.4T A95B | 42% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 19 | Claude Opus 4.7 | 42% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | GLM 5.3 | 42% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 21 | Claude Sonnet 5 | 41% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | GLM 5.2 | 41% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 23 | DeepSeek V4 Pro 0423 | 41% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | Gemini 3.6 Flash | 41% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 25 | Muse Spark | 41% | n/a | n/a | ||||||||||||||||||||||||||||||
| 26 | Qwen3.7 Max | 41% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 27 | Motif 3 (Beta) | Motif-technologies | 40% | n/a | n/a | |||||||||||||||||||||||||||||
| 28 | Claude Opus 4.6 | 40%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 29 | GLM 5.3 Flash | 40% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 30 | Gemini 3 Pro | 40% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | GPT-5.6 Luna | 39% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 32 | MiniMax M3 | 39% | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 33 | DeepSeek V4 Flash 0423 | 39% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 34 | Grok Build 0.1 0616 | Spacexai | 38% | n/a | n/a | |||||||||||||||||||||||||||||
| 35 | Qwen3.8-Flash-Next | 38% | n/a | n/a | ||||||||||||||||||||||||||||||
| 36 | GPT-5.2 | 38% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | Agnes 2.5 Pro Beta | Sapiens-ai | 38% | n/a | n/a | |||||||||||||||||||||||||||||
| 38 | DeepSeek V4 Pro (max) | 38% | n/a | n/a | ||||||||||||||||||||||||||||||
| 39 | Kimi K2.6 | 37% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 40 | Grok 4.3 | 37% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 41 | Motif 3 | Motif-technologies | 37% | n/a | n/a | |||||||||||||||||||||||||||||
| 42 | Gemini 3 Flash | 37%reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 43 | GPT-5.2-Codex | 36% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 44 | MiMo-V2.5-Pro | 36% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Qwen3.7 Plus | 36% | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
| 46 | Kimi K2.7 Code | 35% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 47 | DeepSeek V4 Flash (max) | 35% | n/a | n/a | ||||||||||||||||||||||||||||||
| 48 | Grok 4.20 | 35% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| 49 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 34% | n/a | n/a | ||||||||||||||||||||||||||||||
| 50 | Qwen3.8 27B | 34% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
Frequently asked questions
What is the Humanity's Last Exam benchmark?
A frontier academic benchmark of 2,500 expert-written questions spanning mathematics, the humanities and the natural sciences, in multiple-choice and short-answer formats. Each question has an unambiguous, verifiable answer that cannot be found by quick internet retrieval, and models are scored on accuracy.
Which AI model scores highest on Humanity's Last Exam?
Claude Fable 5 (Anthropic) leads with 55%, from $10.00 per 1M input tokens.
How many models are ranked on Humanity's Last Exam?
430 models carry a Humanity's Last Exam score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI