SWE-bench Verified benchmark

39 AI models ranked on SWE-bench Verified. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Macaron V1 Venti (Other) leads at 86%, from $1.50 per 1M input.

Data via SWE-bench

Models ranked

39

on this benchmark

Top score

86%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs SWE-bench Verified

36 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on SWE-bench Verified

highest score first · 39 models
Maker
#
1Macaron V1 VentiOther86%$1.50$4.50
2DeepSeek V4 Pro 0423DeepSeek81%$0.43$0.87
3MiniMax M3MiniMax81%$0.23$0.96
4Kimi K2.6Moonshot80%$0.65$3.40
5Inkling SmallOther80%$0.45$1.20
6DeepSeek V4 Flash 0423DeepSeek79%$0.09$0.17
7MiMo-V2.5-ProXiaomi79%$0.43$0.87
8Hy3Other78%$0.13$0.53
9GLM 5Zhipu78%$0.60$1.92
10InklingOther78%$0.95$4.05
11Mistral Medium 3.5 128BMistral78%$1.50$7.50
12Qwen3.6 27BAlibaba77%$0.15$0.50
13Qwen3.5 397B A17BAlibaba76%$0.39$2.34
14Muse Glimmer 30BOther76%$0.30$1.20
15MiniMax M2.5MiniMax76%$0.27$1.08
16Macaron V1 TallOther75%$0.45$2.60
17Laguna M.1Other75%$0.20$0.40
18Step 3.5 FlashOther74%$0.10$0.30
19Hy3 previewOther74%$0.18$0.60
20MiniMax M2.1MiniMax74%$0.30$1.20
21Ring-2.6-1TOther74%$0.07$0.63
22GLM 4.7Zhipu74%$0.40$1.75
23Qwen3.6 35B A3BAlibaba73%$0.10$0.45
24Qwen3.5-27BAlibaba72%$0.20$1.56
25Ling-2.6-1TOther72%$0.07$0.63
26Qwen3.5-122B-A10BAlibaba72%$0.25$1.75
27Nvidia Nemotron 3 Ultra 550B A55bNVIDIA72%$0.50$2.20
28Kimi K2 ThinkingMoonshot71%$0.60$1.20
29Laguna XS 2.1Other71%$0.06$0.12
30Kimi K2.5Moonshot71%$0.45$2.25
31Qwen3 Coder NextAlibaba71%$0.12$0.80
32Solar Open2 250BUpstage70%n/an/a
33MiniMax M2MiniMax69%$0.26$1.02
34North Mini CodeCohere68%n/an/a
35gpt-oss-120bOpenAI62%$0.03$0.17
36Ling-2.6-flashOther61%$0.01$0.03
37gpt-oss-20bOpenAI61%$0.01$0.07
38Nemotron 3 120B A12bNVIDIA60%$0.50$1.50
39GLM 4.7 FlashZhipu59%$0.00$0.00
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the SWE-bench Verified benchmark?

SWE-bench Verified is a 500-task subset of SWE-bench, human-validated by OpenAI to remove unsolvable or under-specified problems. Each task is a real GitHub issue from a popular Python repository, and the model must produce a patch that makes the project's hidden test suite pass. Scored as the percentage of issues resolved.

Which AI model scores highest on SWE-bench Verified?

Macaron V1 Venti (Other) leads with 86%, from $1.50 per 1M input tokens.

How many models are ranked on SWE-bench Verified?

39 models carry a SWE-bench Verified score. Scores come from Hugging Face leaderboards; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Hugging Face leaderboards · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI