SciCode benchmark

428 AI models ranked on SciCode. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Claude Fable 5 (Anthropic) leads at 60%, from $10.00 per 1M input.

Data via arXiv:2407.13168

Models ranked

428

on this benchmark

Top score

60%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs SciCode

254 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on SciCode

highest score first · 428 models
Maker
#
1Claude Fable 5Anthropic60%$10.00$50.00
2Gemini 3.1 Pro PreviewGoogle59%$2.00$12.00
3Kimi K3Moonshot59%$2.85$14.25
4Muse Spark 1.1Other58%$1.25$4.25
5GPT-5.6 SolOpenAI57%high$2.00$10.00
6Gemini 3.7 FlashGoogle57%$0.38$1.88
7GPT-5.4OpenAI57%$2.50$15.00
8GLM 5.3Zhipu56%$1.40$4.40
9Muse Spark 1.2Other56%$1.25$4.25
10GPT-5.5OpenAI56%$5.00$30.00
11Gemini 3 ProGoogle56%$2.00$12.00
12Claude Opus 5Anthropic56%$5.00$25.00
13GPT-5.2-CodexOpenAI55%$1.75$14.00
14Claude Opus 4.7Anthropic55%$5.00$25.00
15Grok 4.5xAI54%$2.00$6.00
16GPT-5.6 TerraOpenAI54%$2.00$12.00
17Claude Sonnet 5Anthropic54%$2.00$10.00
18Grok 4.6xAI54%$2.00$6.00
19Kimi K2.6Moonshot53%$0.65$3.40
20Claude Opus 4.8Anthropic53%$5.00$25.00
21GPT-5.3-CodexOpenAI53%$1.75$14.00
22Gemini 3.5 FlashGoogle53%$1.50$9.00
23Qwen3.8 MaxAlibaba53%$1.65$4.95
24Gemini 3.6 FlashGoogle53%$0.75$3.75
25GPT-5.6 LunaOpenAI53%$0.20$1.20
26GPT-5.2OpenAI52%$1.75$14.00
27Claude Opus 4.6Anthropic52%reasoning: adaptive$5.00$25.00
28Qwen3.8 2.4T A95BAlibaba52%$2.00$6.00
29Muse SparkMeta52%n/an/a
30Gemini 3 FlashGoogle51%reasoning: true$0.63$3.75
31GLM 5.2Zhipu50%$0.61$1.98
32GPT-5.5 Instant (May 2026)OpenAI50%n/an/a
33MiMo-V2.5-ProXiaomi50%$0.43$0.87
34Grok Build 0.1 0616Spacexai50%n/an/a
35DeepSeek V4 Pro (max)DeepSeek50%n/an/a
36DeepSeek V4 Flash 0423DeepSeek50%$0.09$0.17
37GPT-5.4 MiniOpenAI50%$0.75$4.50
38Claude Opus 4.5Anthropic50%reasoning: true$5.00$25.00
39DeepSeek V4 Pro 0423DeepSeek49%$0.43$0.87
40Kimi K2.5Moonshot49%$0.45$2.25
41Qwen3.7 MaxAlibaba49%$1.25$3.75
42Inkling SmallOther49%$0.45$1.20
43GPT-5.5 Instant (June 2026)OpenAI49%n/an/a
44Hy3Other48%$0.13$0.53
45Agnes 2.5 Pro BetaSapiens-ai48%n/an/a
46Kimi K2.7 CodeMoonshot47%$0.66$3.40
47Grok 4.3xAI47%$1.25$2.50
48MiniMax M2.7MiniMax47%$0.25$0.55
49Claude Sonnet 4.6Anthropic47%$3.00$15.00
50GPT-5.4 NanoOpenAI47%$0.20$1.25
150 of 428
Page 1 of 9
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the SciCode benchmark?

A research coding benchmark curated by scientists across 16 natural-science sub-fields, where models write Python to solve real scientific problems. Its 80 main problems decompose into 338 subproblems, graded against scientist-annotated gold solutions and test cases.

Which AI model scores highest on SciCode?

Claude Fable 5 (Anthropic) leads with 60%, from $10.00 per 1M input tokens.

How many models are ranked on SciCode?

428 models carry a SciCode score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI