Omniscience benchmark
347 AI models ranked on Omniscience. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Claude Fable 5 (Anthropic) leads at 43.3, from $10.00 per 1M input.
Data via arXiv:2511.13029 →Models ranked
347
on this benchmark
Top score
43.3
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs Omniscience
212 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on Omniscience
highest score first · 347 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 | 43.3 | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 2 | Claude Opus 5 | 37.1 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 3 | Gemini 3.1 Pro Preview | 31.9 | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 4 | Grok 4.6 | 30.5 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 5 | Claude Opus 4.8 | 28.8 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 6 | Muse Spark 1.1 | Other | 28.1 | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 7 | Claude Opus 4.7 | 27.3 | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 8 | Muse Spark 1.2 | Other | 27.2 | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 9 | Gemini 3.7 Flash | 26.5 | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 10 | Grok 4.5 | 25.3 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 11 | Gemini 3.6 Flash | 22.1 | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 12 | GPT-5.6 Sol | 22.0 | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 13 | Gemini 3.5 Flash | 21.2 | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | GPT-5.5 | 20.5 | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 15 | Kimi K3 | 19.7 | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 16 | Grok 4.3 | 18.0 | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 17 | Claude Sonnet 5 | 16.4 | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 18 | Gemini 3 Pro | 15.3 | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 19 | Grok 4.20 | 14.8 | $1.25 | $2.50 | ||||||||||||||||||||||||||||||
| 20 | GLM 5.3 | 14.3 | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 21 | Claude Opus 4.5 | 14reasoning: true | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 22 | Claude Opus 4.6 | 13.7reasoning: adaptive | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 23 | Qwen3.7 Max | 13.5 | $1.25 | $3.75 | ||||||||||||||||||||||||||||||
| 24 | Grok 4.20 0309 | Spacexai | 12.9 | n/a | n/a | |||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 25 | Claude Sonnet 4.6 | 12.2reasoning: adaptive | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 26 | GPT-5.5 Instant (May 2026) | 11.2 | n/a | n/a | ||||||||||||||||||||||||||||||
| 27 | GPT-5.3-Codex | 10.9 | $1.75 | $14.00 | ||||||||||||||||||||||||||||||
| 28 | Motif 3 | Motif-technologies | 10.2 | n/a | n/a | |||||||||||||||||||||||||||||
| 29 | Gemini 3 Flash | 10.1reasoning: true | $0.63 | $3.75 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 30 | Qwen3.6 Max Preview | 9.2 | n/a | n/a | ||||||||||||||||||||||||||||||
| 31 | GLM 5.3 Flash | 7.5 | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 32 | Muse Spark | 7.2 | n/a | n/a | ||||||||||||||||||||||||||||||
| 33 | Grok Build 0.1 0616 | Spacexai | 6.4 | n/a | n/a | |||||||||||||||||||||||||||||
| 34 | GPT-5.4 | 5.8 | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 35 | GPT-5.1 | 5.4 | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 36 | Kimi K2.6 | 5.3 | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 37 | Gemini 3.5 Flash Lite | 5.2 | $0.30 | $2.50 | ||||||||||||||||||||||||||||||
| 38 | MiMo-V2-Pro | 4.6 | n/a | n/a | ||||||||||||||||||||||||||||||
| 39 | GLM 5.2 | 4.4 | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 40 | Qwen3.8 2.4T A95B | 4.3 | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 41 | G9v3-39A5B | Ai9stars | 3.8 | n/a | n/a | |||||||||||||||||||||||||||||
| 42 | GPT-5.5 Instant (June 2026) | 3.7 | n/a | n/a | ||||||||||||||||||||||||||||||
| 43 | Qwen3.8 Max | 3.4 | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 44 | MiMo-V2.5-Pro | 3.3 | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 45 | Grok 4 | 2.1 | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| 46 | Inkling | Other | 2 | $0.95 | $4.05 | |||||||||||||||||||||||||||||
| 47 | MiniMax M3 | 1.4 | $0.23 | $0.96 | ||||||||||||||||||||||||||||||
| 48 | Qwen3.7 Plus | 1.1 | $0.32 | $1.28 | ||||||||||||||||||||||||||||||
| 49 | Qwen3.6 Plus | 0.9 | $0.33 | $1.95 | ||||||||||||||||||||||||||||||
| 50 | GLM 5.1 | 0.8 | $1.05 | $3.50 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
Frequently asked questions
What is the Omniscience benchmark?
AA-Omniscience (Artificial Analysis) measures factual-knowledge reliability and calibration across 6,000 questions in six domains. Its Omniscience Index rewards correct answers and appropriate abstentions while penalising confident wrong answers, so a model that guesses badly can score below zero.
Which AI model scores highest on Omniscience?
Claude Fable 5 (Anthropic) leads with 43.3, from $10.00 per 1M input tokens.
How many models are ranked on Omniscience?
347 models carry a Omniscience score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI