IT-Bench SRE benchmark

35 AI models ranked on IT-Bench SRE. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GPT-5.6 Sol (OpenAI) leads at 56%, from $2.00 per 1M input.

Data via arXiv:2502.05352

Models ranked

35

on this benchmark

Top score

56%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs IT-Bench SRE

33 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on IT-Bench SRE

highest score first · 35 models
Maker
#
1GPT-5.6 SolOpenAI56%$2.00$10.00
2GPT-5.6 TerraOpenAI51%$2.00$12.00
3Kimi K3Moonshot48%$2.85$14.25
4Claude Opus 4.7Anthropic47%$5.00$25.00
5GPT-5.5OpenAI46%$5.00$30.00
6GLM 5.2Zhipu43%$0.61$1.98
7Qwen3.7 MaxAlibaba42%$1.25$3.75
8Gemini 3.5 FlashGoogle40%$1.50$9.00
9GPT-5.6 LunaOpenAI40%$0.20$1.20
10GLM 5.1Zhipu40%$1.05$3.50
11Claude Sonnet 4.6Anthropic40%reasoning: adaptive$3.00$15.00
12DeepSeek V4 Pro 0423DeepSeek38%$0.43$0.87
13DeepSeek V4 Pro (max)DeepSeek38%n/an/a
14MiMo-V2.5-ProXiaomi38%$0.43$0.87
15Gemma 4 31BGoogle37%$0.14$0.40
16Qwen3.5-27BAlibaba35%$0.20$1.56
17GPT-5.4 MiniOpenAI35%$0.75$4.50
18GPT-5.4OpenAI35%$2.50$15.00
19Qwen3.5 397B A17BAlibaba34%$0.39$2.34
20Grok 4.3xAI33%$1.25$2.50
21DeepSeek V4 Flash 0423DeepSeek32%$0.09$0.17
22DeepSeek V4 Flash (max)DeepSeek32%n/an/a
23Kimi K2.6Moonshot31%$0.65$3.40
24Gemini 3.1 Pro PreviewGoogle30%$2.00$12.00
25Step 3.7 FlashOther30%$0.20$1.15
26Claude 4.5 HaikuAnthropic27%reasoning: true$1.00$5.00
27MiniMax M2.7MiniMax26%$0.25$0.55
28GPT-5.4 NanoOpenAI24%$0.20$1.25
29Gemma 4 26B A4bGoogle24%$0.13$0.40
30Qwen3.5-35B-A3BAlibaba22%$0.14$1.00
31Grok 4.1 FastxAI18%$0.20$0.50
32gpt-oss-120bOpenAI6%$0.03$0.17
33Nvidia Nemotron 3 Super 120B A12bNVIDIA1%$0.09$0.40
34Nemotron 3 120B A12bNVIDIA1%$0.50$1.50
35Llama 3.3 70B InstructMeta1%$0.12$0.20
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the IT-Bench SRE benchmark?

IBM's IT-Bench benchmarks AI agents on real-world IT automation; its SRE (Site Reliability Engineering) subset presents Kubernetes-based incident scenarios that must be diagnosed and fully resolved, scored as a binary pass rate. State-of-the-art agents resolved only about 14% in the original paper.

Which AI model scores highest on IT-Bench SRE?

GPT-5.6 Sol (OpenAI) leads with 56%, from $2.00 per 1M input tokens.

How many models are ranked on IT-Bench SRE?

35 models carry a IT-Bench SRE score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI