TerminalBench Hard benchmark

308 AI models ranked on TerminalBench Hard. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GPT-5.6 Sol (OpenAI) leads at 66%, from $2.00 per 1M input.

Data via GitHub: laude-institute/terminal-bench

Models ranked

308

on this benchmark

Top score

66%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs TerminalBench Hard

189 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on TerminalBench Hard

highest score first · 308 models
Maker
#
1GPT-5.6 SolOpenAI66%$2.00$10.00
2Claude Fable 5Anthropic63%$10.00$50.00
3GPT-5.6 TerraOpenAI63%xhigh$2.00$12.00
4GPT-5.5OpenAI61%$5.00$30.00
5Claude Opus 4.8Anthropic58%$5.00$25.00
6GPT-5.4OpenAI58%$2.50$15.00
7Claude Opus 4.7Anthropic55%reasoning: false$5.00$25.00
8Gemini 3.1 Pro PreviewGoogle54%$2.00$12.00
9Claude Sonnet 4.6Anthropic53%reasoning: adaptive$3.00$15.00
10GPT-5.3-CodexOpenAI53%$1.75$14.00
11GPT-5.4 MiniOpenAI52%$0.75$4.50
12GLM 5.2Zhipu51%$0.61$1.98
13Qwen3.7 MaxAlibaba51%$1.25$3.75
14KAT-Coder-Pro V2Other49%$0.30$1.20
15Claude Opus 4.6Anthropic48%$5.00$25.00
16Claude Opus 4.5Anthropic47%reasoning: true$5.00$25.00
17GPT-5.2OpenAI47%$1.75$14.00
18Qwen3.7 PlusAlibaba47%$0.32$1.28
19DeepSeek V4 Pro 0423DeepSeek46%$0.43$0.87
20Gemini 3.5 FlashGoogle46%minimal$1.50$9.00
21DeepSeek V4 Pro (max)DeepSeek46%n/an/a
22GPT-5.1OpenAI45%$1.25$10.00
23Muse SparkMeta45%n/an/a
24Kimi K2.7 CodeMoonshot45%$0.66$3.40
25Kimi K2.6Moonshot44%$0.65$3.40
26Qwen3.6 PlusAlibaba44%$0.33$1.95
27Qwen3.6 Max PreviewAlibaba44%n/an/a
28GLM 5Zhipu43%$0.60$1.92
29GLM 5.1Zhipu43%$1.05$3.50
30MiMo-V2.5-ProXiaomi43%$0.43$0.87
31MiniMax M3MiniMax42%$0.23$0.96
32GPT-5.4 NanoOpenAI42%$0.20$1.25
33GPT-5.5 Instant (May 2026)OpenAI42%n/an/a
34Gemini 3 ProGoogle42%$2.00$12.00
35MiMo-V2.5Xiaomi42%n/an/a
36Qwen3.5 397B A17BAlibaba41%$0.39$2.34
37Grok 4.20 0309Spacexai41%n/an/a
38MiMo-V2-ProXiaomi41%n/an/a
39MiniMax M2.7MiniMax39%$0.25$0.55
40DeepSeek V4 Flash 0423DeepSeek39%high$0.09$0.17
41Gemini 3 FlashGoogle39%reasoning: true$0.63$3.75
42GPT-5OpenAI38%medium$1.25$10.00
43Grok 4xAI38%$3.00$15.00
44Grok 4.3xAI38%$1.25$2.50
45GPT 5 CodexOpenAI38%$1.25$10.00
46Grok 4.20xAI38%$1.25$2.50
47o3OpenAI37%$2.00$8.00
48GPT-5.2-CodexOpenAI37%$1.75$14.00
49Nvidia Nemotron 3 Ultra 550B A55bNVIDIA36%$0.50$2.20
50Gemma 4 31BGoogle36%$0.14$0.40
150 of 308
Page 1 of 7
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the TerminalBench Hard benchmark?

The hardest subset of the Terminal-Bench agentic benchmark, made up of the tasks frontier models complete least often across software engineering, sysadmin, data processing, ML and security. An agent is given an English task and a sandboxed container terminal, and is scored by the pass rate of automated verification tests.

Which AI model scores highest on TerminalBench Hard?

GPT-5.6 Sol (OpenAI) leads with 66%, from $2.00 per 1M input tokens.

How many models are ranked on TerminalBench Hard?

308 models carry a TerminalBench Hard score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI