SWE-bench Pro benchmark

33 AI models ranked on SWE-bench Pro. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Qwen3.8 2.4T A95B (Alibaba) leads at 68%, from $2.00 per 1M input.

Data via Scale AI SWE-bench Pro

Models ranked

33

on this benchmark

Top score

68%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs SWE-bench Pro

31 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on SWE-bench Pro

highest score first · 33 models
Maker
#
1Qwen3.8 2.4T A95BAlibaba68%$2.00$6.00
2Hy4 previewOther66%$0.83$2.50
3Qwen3.8-Flash-NextAlibaba63%n/an/a
4GLM 5.2Zhipu62%$0.61$1.98
5Qwen3.8 27BAlibaba62%$0.40$2.55
6Laguna S 2.1Other59%$0.09$0.18
7MiniMax M3MiniMax59%$0.23$0.96
8Kimi K2.6Moonshot59%$0.65$3.40
9GLM 5p1Zhipu58%$1.40$4.40
10Hy3Other58%$0.13$0.53
11MiMo-V2.5-ProXiaomi57%$0.43$0.87
12Ling-3.0-flashOther57%$0.02$0.06
13Step 3.7 FlashOther56%$0.20$1.15
14MiniMax M2.7MiniMax56%$0.25$0.55
15MiMo-V2.5Xiaomi56%$0.14$0.28
16Inkling SmallOther56%$0.45$1.20
17DeepSeek V4 Pro 0423DeepSeek55%$0.43$0.87
18MiniMax M2.5MiniMax55%$0.27$1.08
19InklingOther54%$0.95$4.05
20Qwen3.6 27BAlibaba54%$0.15$0.50
21Muse Glimmer 30BOther51%$0.30$1.20
22Kimi K2.5Moonshot51%$0.45$2.25
23Qwen3.6 35B A3BAlibaba50%$0.10$0.45
24Laguna M.1Other49%$0.20$0.40
25Laguna XS 2.1Other48%$0.06$0.12
26Qwen3 Coder NextAlibaba44%$0.12$0.80
27North Mini CodeCohere40%n/an/a
28Qwen3 Coder 480B A35b InstructAlibaba39%$0.38$1.50
29MiniMax M2.1MiniMax37%$0.30$1.20
30Kimi K2 InstructMoonshot28%$0.50$2.00
31Qwen3 235B A22BAlibaba21%$0.18$0.54
32gpt-oss-120bOpenAI16%$0.03$0.17
33GLM 4.6Zhipu10%$0.43$1.75
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the SWE-bench Pro benchmark?

SWE-bench Pro (Scale AI) is a harder, contamination-resistant successor to SWE-bench, built from larger and more recent real-world software tasks across many repositories. The model must generate a patch that passes the project's tests. Scored as the percentage of tasks resolved, with frontier models still well short of saturation.

Which AI model scores highest on SWE-bench Pro?

Qwen3.8 2.4T A95B (Alibaba) leads with 68%, from $2.00 per 1M input tokens.

How many models are ranked on SWE-bench Pro?

33 models carry a SWE-bench Pro score. Scores come from Hugging Face leaderboards; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Hugging Face leaderboards · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI