TerminalBench v2.1 benchmark

165 AI models ranked on TerminalBench v2.1. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GPT-5.6 Sol (OpenAI) leads at 90%, from $2.00 per 1M input.

Data via GitHub: harbor-framework/terminal-bench-2-1

Models ranked

165

on this benchmark

Top score

90%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs TerminalBench v2.1

109 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on TerminalBench v2.1

highest score first · 165 models
Maker
#
1GPT-5.6 SolOpenAI90%xhigh$2.00$10.00
2Claude Opus 5Anthropic89%$5.00$25.00
3Grok 4.6xAI88%$2.00$6.00
4GPT-5.6 TerraOpenAI88%$2.00$12.00
5Qwen3.8-Flash-NextAlibaba86%n/an/a
6Gemini 3.7 FlashGoogle86%$0.38$1.88
7Kimi K3Moonshot85%$2.85$14.25
8Claude Opus 4.8Anthropic85%$5.00$25.00
9Claude Fable 5Anthropic85%$10.00$50.00
10GPT-5.5OpenAI84%$5.00$30.00
11GLM 5.3 FlashZhipu84%$0.07$0.25
12GLM 5.3Zhipu84%$1.40$4.40
13Claude Opus 4.7Anthropic83%$5.00$25.00
14Qwen3.8 2.4T A95BAlibaba82%$2.00$6.00
15Grok 4.5xAI82%$2.00$6.00
16Qwen3.8 MaxAlibaba81%$1.65$4.95
17GPT-5.6 LunaOpenAI81%$0.20$1.20
18Claude Sonnet 5Anthropic81%$2.00$10.00
19Muse Spark 1.2Other80%$1.25$4.25
20Qwen3.8 27BAlibaba80%$0.40$2.55
21DeepSeek V4 Flash 0423DeepSeek79%$0.09$0.17
22DeepSeek V4 Pro 0423DeepSeek79%$0.43$0.87
23Gemini 3.5 FlashGoogle79%$1.50$9.00
24GPT-5.4OpenAI78%$2.50$15.00
25GLM 5.2Zhipu78%$0.61$1.98
26Muse Spark 1.1Other78%$1.25$4.25
27Gemini 3.6 FlashGoogle78%$0.75$3.75
28Motif 3Motif-technologies75%n/an/a
29Qwen3.7 MaxAlibaba75%$1.25$3.75
30DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek74%n/an/a
31Gemini 3.1 Pro PreviewGoogle74%$2.00$12.00
32Claude Sonnet 4.6Anthropic71%reasoning: adaptive$3.00$15.00
33Motif 3 (Beta)Motif-technologies71%n/an/a
34KAT-Coder-Pro V2Other70%$0.30$1.20
35Agnes 2.5 Pro BetaSapiens-ai70%n/an/a
36Nex-N2-ProOther68%$0.25$1.00
37Kimi K2.7 CodeMoonshot67%$0.66$3.40
38Agnes 2.5 Pro AlphaSapiens-ai67%n/an/a
39Kimi K2.6Moonshot66%$0.65$3.40
40MiniMax M3MiniMax65%$0.23$0.96
41MiMo-V2.5-ProXiaomi65%$0.43$0.87
42Hy3Other64%$0.13$0.53
43DeepSeek V4 Pro (max)DeepSeek64%n/an/a
44MiMo-V2.5Xiaomi64%n/an/a
45Muse SparkMeta62%n/an/a
46GLM 5.1Zhipu62%$1.05$3.50
47Mimo V2 FlashXiaomi62%$0.10$0.30
48DeepSeek V4 Flash (max)DeepSeek62%n/an/a
49Qwen3.6 PlusAlibaba61%$0.33$1.95
50Qwen3.7 PlusAlibaba61%$0.32$1.28
150 of 165
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the TerminalBench v2.1 benchmark?

A verified release of the Terminal-Bench agentic benchmark: 89 curated real-world tasks (model training, sysadmin, data processing, security, software engineering) an agent must complete in a sandboxed container terminal. It refines v2.0 for bugs, timeouts and reward-hacking robustness, scored by automated verification tests.

Which AI model scores highest on TerminalBench v2.1?

GPT-5.6 Sol (OpenAI) leads with 90%, from $2.00 per 1M input tokens.

How many models are ranked on TerminalBench v2.1?

165 models carry a TerminalBench v2.1 score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI