Long-Context Reasoning benchmark

364 AI models ranked on Long-Context Reasoning. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Muse Spark 1.2 (Other) leads at 83%, from $1.25 per 1M input.

Data via Artificial Analysis AA-LCR

Models ranked

364

on this benchmark

Top score

83%

current leader

Type

Composite

index

Updated

29 Aug 2026

last source fetch

Cost vs Long-Context Reasoning

220 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on Long-Context Reasoning

highest score first · 364 models
Maker
#
1Muse Spark 1.2Other83%$1.25$4.25
2Kimi K3Moonshot83%$2.85$14.25
3Muse Spark 1.1Other81%$1.25$4.25
4Gemini 3.5 FlashGoogle81%$1.50$9.00
5MiniMax M3MiniMax80%$0.23$0.96
6Gemini 3.7 FlashGoogle80%$0.38$1.88
7Muse GlimmerMeta80%n/an/a
8GPT-5.6 TerraOpenAI80%$2.00$12.00
9GPT-5.2OpenAI79%$1.75$14.00
10GPT-5.2-CodexOpenAI79%$1.75$14.00
11Gemini 3.6 FlashGoogle79%$0.75$3.75
12Gemini 3.1 Pro PreviewGoogle79%$2.00$12.00
13GPT-5.5OpenAI79%$5.00$30.00
14GPT-5.6 LunaOpenAI78%$0.20$1.20
15GPT-5.3-CodexOpenAI78%$1.75$14.00
16GLM 5.3 FlashZhipu78%$0.07$0.25
17Agnes 2.5 Pro BetaSapiens-ai78%n/an/a
18DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek78%n/an/a
19GPT-5.4OpenAI78%$2.50$15.00
20MiMo-V2.5-ProXiaomi78%$0.43$0.87
21GPT-5.6 SolOpenAI78%$2.00$10.00
22Qwen3.8 27BAlibaba77%$0.40$2.55
23Claude Sonnet 5Anthropic77%$2.00$10.00
24Muse SparkMeta77%n/an/a
25Qwen3.8-Flash-NextAlibaba77%n/an/a
26GLM 5.2Zhipu77%$0.61$1.98
27Kimi K2.6Moonshot77%$0.65$3.40
28Claude Fable 5Anthropic77%$10.00$50.00
29GPT-5.1OpenAI77%$1.25$10.00
30GPT-5OpenAI76%$1.25$10.00
31GLM 5.3Zhipu76%$1.40$4.40
32Nex-N2-ProOther76%$0.25$1.00
33Claude Opus 4.5Anthropic76%reasoning: true$5.00$25.00
34Claude Opus 5Anthropic76%$5.00$25.00
35DeepSeek V4 Pro 0423DeepSeek75%$0.43$0.87
36Claude Opus 4.7Anthropic75%$5.00$25.00
37MiniMax M2.7MiniMax75%$0.25$0.55
38Qwen3.8 2.4T A95BAlibaba75%$2.00$6.00
39Kimi K2.7 CodeMoonshot75%$0.66$3.40
40Grok 4.6xAI75%$2.00$6.00
41Qwen3.7 MaxAlibaba75%$1.25$3.75
42Hy3Other75%$0.13$0.53
43Gemini 3.5 Flash LiteGoogle75%$0.30$2.50
44DeepSeek V4 Flash 0423DeepSeek74%$0.09$0.17
45Claude Opus 4.6Anthropic74%reasoning: adaptive$5.00$25.00
46Qwen3.8 MaxAlibaba74%$1.65$4.95
47Claude Sonnet 4.6Anthropic74%reasoning: adaptive$3.00$15.00
48Grok 4.5xAI74%$2.00$6.00
49Claude 4.5 HaikuAnthropic74%reasoning: true$1.00$5.00
50Qwen3.6 27BAlibaba73%$0.15$0.50
150 of 364
Page 1 of 8
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the Long-Context Reasoning benchmark?

Measures a model’s ability to retrieve, synthesise and reason across very long inputs spanning tens to hundreds of thousands of tokens, beyond simple fact lookup. Artificial Analysis reports this as AA-LCR; the broader category also draws on evaluations such as RULER and LongBench.

Which AI model scores highest on Long-Context Reasoning?

Muse Spark 1.2 (Other) leads with 83%, from $1.25 per 1M input tokens.

How many models are ranked on Long-Context Reasoning?

364 models carry a Long-Context Reasoning score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI