TerminalBench v2.1 benchmark
165 AI models ranked on TerminalBench v2.1. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
GPT-5.6 Sol (OpenAI) leads at 90%, from $2.00 per 1M input.
Data via GitHub: harbor-framework/terminal-bench-2-1 →Models ranked
165
on this benchmark
Top score
90%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs TerminalBench v2.1
109 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on TerminalBench v2.1
highest score first · 165 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 90%xhigh | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 2 | Claude Opus 5 | 89% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 3 | Grok 4.6 | 88% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 4 | GPT-5.6 Terra | 88% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 5 | Qwen3.8-Flash-Next | 86% | n/a | n/a | ||||||||||||||||||||||||||||||
| 6 | Gemini 3.7 Flash | 86% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 7 | Kimi K3 | 85% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 8 | Claude Opus 4.8 | 85% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 9 | Claude Fable 5 | 85% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 10 | GPT-5.5 | 84% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 11 | GLM 5.3 Flash | 84% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 12 | GLM 5.3 | 84% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 13 | Claude Opus 4.7 | 83% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 14 | Qwen3.8 2.4T A95B | 82% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 15 | Grok 4.5 | 82% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 16 | Qwen3.8 Max | 81% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 17 | GPT-5.6 Luna | 81% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 18 | Claude Sonnet 5 | 81% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 19 | Muse Spark 1.2 | Other | 80% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 20 | Qwen3.8 27B | 80% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 21 | DeepSeek V4 Flash 0423 | 79% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | DeepSeek V4 Pro 0423 | 79% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 23 | Gemini 3.5 Flash | 79% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| 24 | GPT-5.4 | 78% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| 25 | GLM 5.2 | 78% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 26 | Muse Spark 1.1 | Other | 78% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 27 | Gemini 3.6 Flash | 78% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 28 | Motif 3 | Motif-technologies | 75% | n/a | n/a | |||||||||||||||||||||||||||||
| 29 | Qwen3.7 Max | 75% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 30 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 74% | n/a | n/a | ||||||||||||||||||||||||||||||
| 31 | Gemini 3.1 Pro Preview | 74% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 32 | Claude Sonnet 4.6 | 71%reasoning: adaptive | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| 33 | Motif 3 (Beta) | Motif-technologies | 71% | n/a | n/a | |||||||||||||||||||||||||||||
| 34 | KAT-Coder-Pro V2 | Other | 70% | $0.30 | $1.20 | |||||||||||||||||||||||||||||
| 35 | Agnes 2.5 Pro Beta | Sapiens-ai | 70% | n/a | n/a | |||||||||||||||||||||||||||||
| 36 | Nex-N2-Pro | Other | 68% | $0.25 | $1.00 | |||||||||||||||||||||||||||||
| 37 | Kimi K2.7 Code | 67% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 38 | Agnes 2.5 Pro Alpha | Sapiens-ai | 67% | n/a | n/a | |||||||||||||||||||||||||||||
| 39 | Kimi K2.6 | 66% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| 40 | MiniMax M3 | 65% | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 41 | MiMo-V2.5-Pro | 65% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| 42 | Hy3 | Other | 64% | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| 43 | DeepSeek V4 Pro (max) | 64% | n/a | n/a | ||||||||||||||||||||||||||||||
| 44 | MiMo-V2.5 | 64% | n/a | n/a | ||||||||||||||||||||||||||||||
| 45 | Muse Spark | 62% | n/a | n/a | ||||||||||||||||||||||||||||||
| 46 | GLM 5.1 | 62% | $1.05 | $3.50 | ||||||||||||||||||||||||||||||
| 47 | Mimo V2 Flash | 62% | $0.10 | $0.30 | ||||||||||||||||||||||||||||||
| 48 | DeepSeek V4 Flash (max) | 62% | n/a | n/a | ||||||||||||||||||||||||||||||
| 49 | Qwen3.6 Plus | 61% | $0.33 | $1.95 | ||||||||||||||||||||||||||||||
| 50 | Qwen3.7 Plus | 61% | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
Frequently asked questions
What is the TerminalBench v2.1 benchmark?
A verified release of the Terminal-Bench agentic benchmark: 89 curated real-world tasks (model training, sysadmin, data processing, security, software engineering) an agent must complete in a sandboxed container terminal. It refines v2.0 for bugs, timeouts and reward-hacking robustness, scored by automated verification tests.
Which AI model scores highest on TerminalBench v2.1?
GPT-5.6 Sol (OpenAI) leads with 90%, from $2.00 per 1M input tokens.
How many models are ranked on TerminalBench v2.1?
165 models carry a TerminalBench v2.1 score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI