SWE-bench Verified benchmark
39 AI models ranked on SWE-bench Verified. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Models ranked
39
on this benchmark
Top score
86%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs SWE-bench Verified
36 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on SWE-bench Verified
highest score first · 39 models| # | ||||||
|---|---|---|---|---|---|---|
| 1 | Macaron V1 Venti | Other | 86% | $1.50 | $4.50 | |
| 2 | DeepSeek V4 Pro 0423 | 81% | $0.43 | $0.87 | ||
| 3 | MiniMax M3 | 81% | $0.23 | $0.96 | ||
| 4 | Kimi K2.6 | 80% | $0.65 | $3.40 | ||
| 5 | Inkling Small | Other | 80% | $0.45 | $1.20 | |
| 6 | DeepSeek V4 Flash 0423 | 79% | $0.09 | $0.17 | ||
| 7 | MiMo-V2.5-Pro | 79% | $0.43 | $0.87 | ||
| 8 | Hy3 | Other | 78% | $0.13 | $0.53 | |
| 9 | GLM 5 | 78% | $0.60 | $1.92 | ||
| 10 | Inkling | Other | 78% | $0.95 | $4.05 | |
| 11 | Mistral Medium 3.5 128B | 78% | $1.50 | $7.50 | ||
| 12 | Qwen3.6 27B | 77% | $0.15 | $0.50 | ||
| 13 | Qwen3.5 397B A17B | 76% | $0.39 | $2.34 | ||
| 14 | Muse Glimmer 30B | Other | 76% | $0.30 | $1.20 | |
| 15 | MiniMax M2.5 | 76% | $0.27 | $1.08 | ||
| 16 | Macaron V1 Tall | Other | 75% | $0.45 | $2.60 | |
| 17 | Laguna M.1 | Other | 75% | $0.20 | $0.40 | |
| 18 | Step 3.5 Flash | Other | 74% | $0.10 | $0.30 | |
| 19 | Hy3 preview | Other | 74% | $0.18 | $0.60 | |
| 20 | MiniMax M2.1 | 74% | $0.30 | $1.20 | ||
| 21 | Ring-2.6-1T | Other | 74% | $0.07 | $0.63 | |
| 22 | GLM 4.7 | 74% | $0.40 | $1.75 | ||
| 23 | Qwen3.6 35B A3B | 73% | $0.10 | $0.45 | ||
| 24 | Qwen3.5-27B | 72% | $0.20 | $1.56 | ||
| 25 | Ling-2.6-1T | Other | 72% | $0.07 | $0.63 | |
| 26 | Qwen3.5-122B-A10B | 72% | $0.25 | $1.75 | ||
| 27 | Nvidia Nemotron 3 Ultra 550B A55b | 72% | $0.50 | $2.20 | ||
| 28 | Kimi K2 Thinking | 71% | $0.60 | $1.20 | ||
| 29 | Laguna XS 2.1 | Other | 71% | $0.06 | $0.12 | |
| 30 | Kimi K2.5 | 71% | $0.45 | $2.25 | ||
| 31 | Qwen3 Coder Next | 71% | $0.12 | $0.80 | ||
| 32 | Solar Open2 250B | Upstage | 70% | n/a | n/a | |
| 33 | MiniMax M2 | 69% | $0.26 | $1.02 | ||
| 34 | North Mini Code | 68% | n/a | n/a | ||
| 35 | gpt-oss-120b | 62% | $0.03 | $0.17 | ||
| 36 | Ling-2.6-flash | Other | 61% | $0.01 | $0.03 | |
| 37 | gpt-oss-20b | 61% | $0.01 | $0.07 | ||
| 38 | Nemotron 3 120B A12b | 60% | $0.50 | $1.50 | ||
| 39 | GLM 4.7 Flash | 59% | $0.00 | $0.00 |
Frequently asked questions
What is the SWE-bench Verified benchmark?
SWE-bench Verified is a 500-task subset of SWE-bench, human-validated by OpenAI to remove unsolvable or under-specified problems. Each task is a real GitHub issue from a popular Python repository, and the model must produce a patch that makes the project's hidden test suite pass. Scored as the percentage of issues resolved.
Which AI model scores highest on SWE-bench Verified?
Macaron V1 Venti (Other) leads with 86%, from $1.50 per 1M input tokens.
How many models are ranked on SWE-bench Verified?
39 models carry a SWE-bench Verified score. Scores come from Hugging Face leaderboards; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI