Coding Index benchmark

184 AI models ranked on Coding Index. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GPT-5.6 Sol (OpenAI) leads at 78.3, from $2.00 per 1M input.

Data via Artificial Analysis

Models ranked

184

on this benchmark

Top score

78.3

current leader

Type

Composite

index

Updated

29 Aug 2026

last source fetch

Cost vs Coding Index

120 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on Coding Index

highest score first · 184 models
Maker
#
1GPT-5.6 SolOpenAI78.3xhigh$2.00$10.00
2Claude Opus 5Anthropic78.0$5.00$25.00
3Grok 4.6xAI76.8$2.00$6.00
4GPT-5.6 TerraOpenAI76.7$2.00$12.00
5Claude Fable 5Anthropic76.5$10.00$50.00
6Kimi K3Moonshot76.2$2.85$14.25
7Gemini 3.7 FlashGoogle76.1$0.38$1.88
8GPT-5.5OpenAI74.9$5.00$30.00
9GLM 5.3Zhipu74.8$1.40$4.40
10Claude Opus 4.8Anthropic74.3$5.00$25.00
11Claude Opus 4.7Anthropic73.6$5.00$25.00
12Qwen3.8-Flash-NextAlibaba73.1n/an/a
13Grok 4.5xAI72.4$2.00$6.00
14Muse Spark 1.2Other72.2$1.25$4.25
15Qwen3.8 2.4T A95BAlibaba71.9$2.00$6.00
16Qwen3.8 MaxAlibaba71.8$1.65$4.95
17Claude Sonnet 5Anthropic71.5$2.00$10.00
18GLM 5.3 FlashZhipu71.5$0.07$0.25
19GPT-5.6 LunaOpenAI71.4$0.20$1.20
20Muse Spark 1.1Other71.3$1.25$4.25
21GPT-5.4OpenAI71.1$2.50$15.00
22Gemini 3.5 FlashGoogle70.1$1.50$9.00
23Gemini 3.6 FlashGoogle69.2$0.75$3.75
24DeepSeek V4 Flash 0423DeepSeek69.1$0.09$0.17
25DeepSeek V4 Pro 0423DeepSeek68.8$0.43$0.87
26Gemini 3.1 Pro PreviewGoogle68.8$2.00$12.00
27GLM 5.2Zhipu68.8$0.61$1.98
28Qwen3.8 27BAlibaba68.1$0.40$2.55
29Qwen3.7 MaxAlibaba66.0$1.25$3.75
30DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek65.0n/an/a
31Motif 3Motif-technologies63.5n/an/a
32Claude Sonnet 4.6Anthropic63.0reasoning: adaptive$3.00$15.00
33Agnes 2.5 Pro BetaSapiens-ai62.3n/an/a
34Motif 3 (Beta)Motif-technologies62.0n/an/a
35Kimi K2.6Moonshot61.8$0.65$3.40
36Kimi K2.7 CodeMoonshot60.8$0.66$3.40
37MiMo-V2.5-ProXiaomi60.2$0.43$0.87
38KAT-Coder-Pro V2Other59.5$0.30$1.20
39DeepSeek V4 Pro (max)DeepSeek59.4n/an/a
40Nex-N2-ProOther59.1$0.25$1.00
41Hy3Other58.8$0.13$0.53
42Agnes 2.5 Pro AlphaSapiens-ai58.8n/an/a
43Muse SparkMeta58.6n/an/a
44MiniMax M3MiniMax58.6$0.23$0.96
45MiMo-V2.5Xiaomi56.8n/an/a
46DeepSeek V4 Flash (max)DeepSeek56.2n/an/a
47GPT-5.4 MiniOpenAI56.1$0.75$4.50
48GPT-5.4 NanoOpenAI56.1$0.20$1.25
49Qwen3.7 PlusAlibaba55.9$0.32$1.28
50GLM 5.1Zhipu55.8$1.05$3.50
150 of 184
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the Coding Index benchmark?

Artificial Analysis's composite coding score (0 to 100), averaging coding evaluations such as Terminal-Bench (terminal-based software-engineering and sysadmin tasks) and SciCode (scientific-computing Python problems) and normalising to a 0 to 100 scale. Higher means stronger overall coding performance.

Which AI model scores highest on Coding Index?

GPT-5.6 Sol (OpenAI) leads with 78.3, from $2.00 per 1M input tokens.

How many models are ranked on Coding Index?

184 models carry a Coding Index score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI