SciCode benchmark
428 AI models ranked on SciCode. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Claude Fable 5 (Anthropic) leads at 60%, from $10.00 per 1M input.
Data via arXiv:2407.13168 →Models ranked
428
on this benchmark
Top score
60%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs SciCode
254 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on SciCode
highest score first · 428 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 60% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 2 | Gemini 3.1 Pro Preview | 59% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 3 | Kimi K3 | 59% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 4 | Muse Spark 1.1 | Other | 58% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 5 | GPT-5.6 Sol | 57%high | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 6 | Gemini 3.7 Flash | 57% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 7 | GPT-5.4 | 57% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 8 | GLM 5.3 | 56% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 9 | Muse Spark 1.2 | Other | 56% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 10 | GPT-5.5 | 56% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 11 | Gemini 3 Pro | 56% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 12 | Claude Opus 5 | 56% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 13 | GPT-5.2-Codex | 55% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 14 | Claude Opus 4.7 | 55% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 15 | Grok 4.5 | 54% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 16 | GPT-5.6 Terra | 54% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 17 | Claude Sonnet 5 | 54% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 18 | Grok 4.6 | 54% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 19 | Kimi K2.6 | 53% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | Claude Opus 4.8 | 53% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 21 | GPT-5.3-Codex | 53% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 22 | Gemini 3.5 Flash | 53% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 23 | Qwen3.8 Max | 53% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 24 | Gemini 3.6 Flash | 53% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 25 | GPT-5.6 Luna | 53% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 26 | GPT-5.2 | 52% | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 27 | Claude Opus 4.6 | 52%reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 28 | Qwen3.8 2.4T A95B | 52% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 29 | Muse Spark | 52% | n/a | n/a | ||||||||||||||||||||||||||||||
| 30 | Gemini 3 Flash | 51%reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | GLM 5.2 | 50% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 32 | GPT-5.5 Instant (May 2026) | 50% | n/a | n/a | ||||||||||||||||||||||||||||||
| 33 | MiMo-V2.5-Pro | 50% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 34 | Grok Build 0.1 0616 | Spacexai | 50% | n/a | n/a | |||||||||||||||||||||||||||||
| 35 | DeepSeek V4 Pro (max) | 50% | n/a | n/a | ||||||||||||||||||||||||||||||
| 36 | DeepSeek V4 Flash 0423 | 50% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | GPT-5.4 Mini | 50% | $0.75 | $4.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 38 | Claude Opus 4.5 | 50%reasoning: true | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 39 | DeepSeek V4 Pro 0423 | 49% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 40 | Kimi K2.5 | 49% | $0.45 | $2.25 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 41 | Qwen3.7 Max | 49% | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 42 | Inkling Small | Other | 49% | $0.45 | $1.20 | |||||||||||||||||||||||||||||
| 43 | GPT-5.5 Instant (June 2026) | 49% | n/a | n/a | ||||||||||||||||||||||||||||||
| 44 | Hy3 | Other | 48% | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Agnes 2.5 Pro Beta | Sapiens-ai | 48% | n/a | n/a | |||||||||||||||||||||||||||||
| 46 | Kimi K2.7 Code | 47% | $0.66 | $3.40 | ||||||||||||||||||||||||||||||
| 47 | Grok 4.3 | 47% | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 48 | MiniMax M2.7 | 47% | $0.25 | $0.55 | ||||||||||||||||||||||||||||||
| 49 | Claude Sonnet 4.6 | 47% | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 50 | GPT-5.4 Nano | 47% | $0.20 | $1.25 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
Frequently asked questions
What is the SciCode benchmark?
A research coding benchmark curated by scientists across 16 natural-science sub-fields, where models write Python to solve real scientific problems. Its 80 main problems decompose into 338 subproblems, graded against scientist-annotated gold solutions and test cases.
Which AI model scores highest on SciCode?
Claude Fable 5 (Anthropic) leads with 60%, from $10.00 per 1M input tokens.
How many models are ranked on SciCode?
428 models carry a SciCode score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI