Long-Context Reasoning benchmark
364 AI models ranked on Long-Context Reasoning. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Muse Spark 1.2 (Other) leads at 83%, from $1.25 per 1M input.
Data via Artificial Analysis AA-LCR →Models ranked
364
on this benchmark
Top score
83%
current leader
Type
Composite
index
Updated
29 Aug 2026
last source fetch
Cost vs Long-Context Reasoning
220 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on Long-Context Reasoning
highest score first · 364 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Muse Spark 1.2 | Other | 83% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 2 | Kimi K3 | 83% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 3 | Muse Spark 1.1 | Other | 81% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 4 | Gemini 3.5 Flash | 81% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 5 | MiniMax M3 | 80% | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 6 | Gemini 3.7 Flash | 80% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 7 | Muse Glimmer | 80% | n/a | n/a | ||||||||||||||||||||||||||||||
| 8 | GPT-5.6 Terra | 80% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 9 | GPT-5.2 | 79% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | GPT-5.2-Codex | 79% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 11 | Gemini 3.6 Flash | 79% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 12 | Gemini 3.1 Pro Preview | 79% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 13 | GPT-5.5 | 79% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | GPT-5.6 Luna | 78% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 15 | GPT-5.3-Codex | 78% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 16 | GLM 5.3 Flash | 78% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 17 | Agnes 2.5 Pro Beta | Sapiens-ai | 78% | n/a | n/a | |||||||||||||||||||||||||||||
| 18 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 78% | n/a | n/a | ||||||||||||||||||||||||||||||
| 19 | GPT-5.4 | 78% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | MiMo-V2.5-Pro | 78% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 21 | GPT-5.6 Sol | 78% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | Qwen3.8 27B | 77% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 23 | Claude Sonnet 5 | 77% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | Muse Spark | 77% | n/a | n/a | ||||||||||||||||||||||||||||||
| 25 | Qwen3.8-Flash-Next | 77% | n/a | n/a | ||||||||||||||||||||||||||||||
| 26 | GLM 5.2 | 77% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 27 | Kimi K2.6 | 77% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 28 | Claude Fable 5 | 77% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 29 | GPT-5.1 | 77% | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 30 | GPT-5 | 76% | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | GLM 5.3 | 76% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 32 | Nex-N2-Pro | Other | 76% | $0.25 | $1.00 | |||||||||||||||||||||||||||||
| 33 | Claude Opus 4.5 | 76%reasoning: true | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 34 | Claude Opus 5 | 76% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 35 | DeepSeek V4 Pro 0423 | 75% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 36 | Claude Opus 4.7 | 75% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | MiniMax M2.7 | 75% | $0.25 | $0.55 | ||||||||||||||||||||||||||||||
| 38 | Qwen3.8 2.4T A95B | 75% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 39 | Kimi K2.7 Code | 75% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 40 | Grok 4.6 | 75% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 41 | Qwen3.7 Max | 75% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 42 | Hy3 | Other | 75% | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 43 | Gemini 3.5 Flash Lite | 75% | $0.30 | $2.50 | ||||||||||||||||||||||||||||||
| 44 | DeepSeek V4 Flash 0423 | 74% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Claude Opus 4.6 | 74%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 46 | Qwen3.8 Max | 74% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 47 | Claude Sonnet 4.6 | 74%reasoning: adaptive | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 48 | Grok 4.5 | 74% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 49 | Claude 4.5 Haiku | 74%reasoning: true | $1.00 | $5.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 50 | Qwen3.6 27B | 73% | $0.15 | $0.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
Frequently asked questions
What is the Long-Context Reasoning benchmark?
Measures a model’s ability to retrieve, synthesise and reason across very long inputs spanning tens to hundreds of thousands of tokens, beyond simple fact lookup. Artificial Analysis reports this as AA-LCR; the broader category also draws on evaluations such as RULER and LongBench.
Which AI model scores highest on Long-Context Reasoning?
Muse Spark 1.2 (Other) leads with 83%, from $1.25 per 1M input tokens.
How many models are ranked on Long-Context Reasoning?
364 models carry a Long-Context Reasoning score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI