Humanity's Last Exam benchmark

430 AI models ranked on Humanity's Last Exam. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Claude Fable 5 (Anthropic) leads at 55%, from $10.00 per 1M input.

Data via arXiv:2501.14249

Models ranked

430

on this benchmark

Top score

55%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs Humanity's Last Exam

256 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on Humanity's Last Exam

highest score first · 430 models
Maker
#
1Claude Fable 5Anthropic55%$10.00$50.00
2Claude Opus 5Anthropic55%$5.00$25.00
3GPT-5.6 SolOpenAI49%$2.00$10.00
4Claude Opus 4.8Anthropic49%$5.00$25.00
5Gemini 3.7 FlashGoogle48%$0.38$1.88
6Gemini 3.1 Pro PreviewGoogle47%$2.00$12.00
7Kimi K3Moonshot47%$2.85$14.25
8Muse Spark 1.1Other46%$1.25$4.25
9GPT-5.5OpenAI46%$5.00$30.00
10Muse Spark 1.2Other45%$1.25$4.25
11GPT-5.4OpenAI44%$2.50$15.00
12Qwen3.8 MaxAlibaba43%$1.65$4.95
13GPT-5.6 TerraOpenAI43%$2.00$12.00
14Grok 4.6xAI43%$2.00$6.00
15Gemini 3.5 FlashGoogle43%$1.50$9.00
16Grok 4.5xAI43%$2.00$6.00
17GPT-5.3-CodexOpenAI42%$1.75$14.00
18Qwen3.8 2.4T A95BAlibaba42%$2.00$6.00
19Claude Opus 4.7Anthropic42%$5.00$25.00
20GLM 5.3Zhipu42%$1.40$4.40
21Claude Sonnet 5Anthropic41%$2.00$10.00
22GLM 5.2Zhipu41%$0.61$1.98
23DeepSeek V4 Pro 0423DeepSeek41%$0.43$0.87
24Gemini 3.6 FlashGoogle41%$0.75$3.75
25Muse SparkMeta41%n/an/a
26Qwen3.7 MaxAlibaba41%$1.25$3.75
27Motif 3 (Beta)Motif-technologies40%n/an/a
28Claude Opus 4.6Anthropic40%reasoning: adaptive$5.00$25.00
29GLM 5.3 FlashZhipu40%$0.07$0.25
30Gemini 3 ProGoogle40%$2.00$12.00
31GPT-5.6 LunaOpenAI39%$0.20$1.20
32MiniMax M3MiniMax39%$0.23$0.96
33DeepSeek V4 Flash 0423DeepSeek39%$0.09$0.17
34Grok Build 0.1 0616Spacexai38%n/an/a
35Qwen3.8-Flash-NextAlibaba38%n/an/a
36GPT-5.2OpenAI38%$1.75$14.00
37Agnes 2.5 Pro BetaSapiens-ai38%n/an/a
38DeepSeek V4 Pro (max)DeepSeek38%n/an/a
39Kimi K2.6Moonshot37%$0.65$3.40
40Grok 4.3xAI37%$1.25$2.50
41Motif 3Motif-technologies37%n/an/a
42Gemini 3 FlashGoogle37%reasoning: true$0.63$3.75
43GPT-5.2-CodexOpenAI36%$1.75$14.00
44MiMo-V2.5-ProXiaomi36%$0.43$0.87
45Qwen3.7 PlusAlibaba36%$0.32$1.28
46Kimi K2.7 CodeMoonshot35%$0.66$3.40
47DeepSeek V4 Flash (max)DeepSeek35%n/an/a
48Grok 4.20xAI35%$1.25$2.50
49DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek34%n/an/a
50Qwen3.8 27BAlibaba34%$0.40$2.55
150 of 430
Page 1 of 9
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the Humanity's Last Exam benchmark?

A frontier academic benchmark of 2,500 expert-written questions spanning mathematics, the humanities and the natural sciences, in multiple-choice and short-answer formats. Each question has an unambiguous, verifiable answer that cannot be found by quick internet retrieval, and models are scored on accuracy.

Which AI model scores highest on Humanity's Last Exam?

Claude Fable 5 (Anthropic) leads with 55%, from $10.00 per 1M input tokens.

How many models are ranked on Humanity's Last Exam?

430 models carry a Humanity's Last Exam score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI