Agentic Index benchmark

159 AI models ranked on Agentic Index. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Claude Opus 5 (Anthropic) leads at 59.2, from $5.00 per 1M input.

Data via Artificial Analysis

Models ranked

159

on this benchmark

Top score

59.2

current leader

Type

Composite

index

Updated

29 Aug 2026

last source fetch

Cost vs Agentic Index

107 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on Agentic Index

highest score first · 159 models
Maker
#
1Claude Opus 5Anthropic59.2$5.00$25.00
2GLM 5.3Zhipu59.1$1.40$4.40
3Grok 4.6xAI58.7$2.00$6.00
4Qwen3.8 MaxAlibaba58.4$1.65$4.95
5GLM 5.3 FlashZhipu58.2$0.07$0.25
6GPT-5.6 SolOpenAI57.8$2.00$10.00
7Qwen3.8 2.4T A95BAlibaba57.1$2.00$6.00
8Claude Fable 5Anthropic56.6$10.00$50.00
9Qwen3.8-Flash-NextAlibaba56.4n/an/a
10Kimi K3Moonshot54.3$2.85$14.25
11DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek52.9n/an/a
12Qwen3.8 27BAlibaba50.9$0.40$2.55
13GPT-5.6 TerraOpenAI50.2$2.00$12.00
14Claude Sonnet 5Anthropic49.7$2.00$10.00
15DeepSeek V4 Pro 0423DeepSeek49.6$0.43$0.87
16Claude Opus 4.8Anthropic49.4$5.00$25.00
17Muse Spark 1.2Other49.3$1.25$4.25
18Grok 4.5xAI48.9$2.00$6.00
19DeepSeek V4 Flash 0423DeepSeek48.4$0.09$0.17
20GPT-5.5OpenAI47.4$5.00$30.00
21GPT-5.6 LunaOpenAI46.9$0.20$1.20
22Claude Opus 4.7Anthropic46.3$5.00$25.00
23GLM 5.2Zhipu45.7$0.61$1.98
24Gemini 3.7 FlashGoogle45.1$0.38$1.88
25GPT-5.4OpenAI44.2$2.50$15.00
26Agnes 2.5 Pro BetaSapiens-ai43.8n/an/a
27Claude Sonnet 4.6Anthropic42.1reasoning: adaptive$3.00$15.00
28Gemini 3.6 FlashGoogle40.5$0.75$3.75
29Muse Spark 1.1Other39.7$1.25$4.25
30Gemini 3.5 FlashGoogle39.7$1.50$9.00
31DeepSeek V4 Pro (max)DeepSeek37.8n/an/a
32Motif 3Motif-technologies37.6n/an/a
33MiniMax M3MiniMax36.1$0.23$0.96
34Motif 3 (Beta)Motif-technologies34.9n/an/a
35JT-4.1 Flash 236B A21BChina-mobile34.9n/an/a
36InklingOther34.1$0.95$4.05
37DeepSeek V4 Flash (max)DeepSeek33.7n/an/a
38Solar Pro 4Other33.6$0.03$0.12
39G9v3-39A5BAi9stars32.9n/an/a
40Inkling SmallOther31.9$0.45$1.20
41GPT-5.4 MiniOpenAI31.5$0.75$4.50
42Hy3Other31.4$0.13$0.53
43Nex-N2-ProOther31.2$0.25$1.00
44Kimi K2.6Moonshot31.2$0.65$3.40
45Qwen3.7 MaxAlibaba30.9$1.25$3.75
46GLM 5.1Zhipu30.6$1.05$3.50
47Kimi K2.7 CodeMoonshot30.3$0.66$3.40
48GPT-5.4 NanoOpenAI29.7$0.20$1.25
49MiMo-V2.5-ProXiaomi29.5$0.43$0.87
50Ling-3.0-flashOther29.3$0.02$0.06
150 of 159
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the Agentic Index benchmark?

Artificial Analysis's composite agentic score (0 to 100) measuring tool use, planning, autonomy and multi-step problem solving, computed by averaging its underlying agentic evaluations and normalising to a 0 to 100 scale.

Which AI model scores highest on Agentic Index?

Claude Opus 5 (Anthropic) leads with 59.2, from $5.00 per 1M input tokens.

How many models are ranked on Agentic Index?

159 models carry a Agentic Index score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI