Omniscience benchmark

347 AI models ranked on Omniscience. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Claude Fable 5 (Anthropic) leads at 43.3, from $10.00 per 1M input.

Data via arXiv:2511.13029

Models ranked

347

on this benchmark

Top score

43.3

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs Omniscience

212 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on Omniscience

highest score first · 347 models
Maker
#
1Claude Fable 5Anthropic43.3$10.00$50.00
2Claude Opus 5Anthropic37.1$5.00$25.00
3Gemini 3.1 Pro PreviewGoogle31.9$2.00$12.00
4Grok 4.6xAI30.5$2.00$6.00
5Claude Opus 4.8Anthropic28.8$5.00$25.00
6Muse Spark 1.1Other28.1$1.25$4.25
7Claude Opus 4.7Anthropic27.3$5.00$25.00
8Muse Spark 1.2Other27.2$1.25$4.25
9Gemini 3.7 FlashGoogle26.5$0.38$1.88
10Grok 4.5xAI25.3$2.00$6.00
11Gemini 3.6 FlashGoogle22.1$0.75$3.75
12GPT-5.6 SolOpenAI22.0$2.00$10.00
13Gemini 3.5 FlashGoogle21.2$1.50$9.00
14GPT-5.5OpenAI20.5$5.00$30.00
15Kimi K3Moonshot19.7$2.85$14.25
16Grok 4.3xAI18.0$1.25$2.50
17Claude Sonnet 5Anthropic16.4$2.00$10.00
18Gemini 3 ProGoogle15.3$2.00$12.00
19Grok 4.20xAI14.8$1.25$2.50
20GLM 5.3Zhipu14.3$1.40$4.40
21Claude Opus 4.5Anthropic14reasoning: true$5.00$25.00
22Claude Opus 4.6Anthropic13.7reasoning: adaptive$5.00$25.00
23Qwen3.7 MaxAlibaba13.5$1.25$3.75
24Grok 4.20 0309Spacexai12.9n/an/a
25Claude Sonnet 4.6Anthropic12.2reasoning: adaptive$3.00$15.00
26GPT-5.5 Instant (May 2026)OpenAI11.2n/an/a
27GPT-5.3-CodexOpenAI10.9$1.75$14.00
28Motif 3Motif-technologies10.2n/an/a
29Gemini 3 FlashGoogle10.1reasoning: true$0.63$3.75
30Qwen3.6 Max PreviewAlibaba9.2n/an/a
31GLM 5.3 FlashZhipu7.5$0.07$0.25
32Muse SparkMeta7.2n/an/a
33Grok Build 0.1 0616Spacexai6.4n/an/a
34GPT-5.4OpenAI5.8$2.50$15.00
35GPT-5.1OpenAI5.4$1.25$10.00
36Kimi K2.6Moonshot5.3$0.65$3.40
37Gemini 3.5 Flash LiteGoogle5.2$0.30$2.50
38MiMo-V2-ProXiaomi4.6n/an/a
39GLM 5.2Zhipu4.4$0.61$1.98
40Qwen3.8 2.4T A95BAlibaba4.3$2.00$6.00
41G9v3-39A5BAi9stars3.8n/an/a
42GPT-5.5 Instant (June 2026)OpenAI3.7n/an/a
43Qwen3.8 MaxAlibaba3.4$1.65$4.95
44MiMo-V2.5-ProXiaomi3.3$0.43$0.87
45Grok 4xAI2.1$3.00$15.00
46InklingOther2$0.95$4.05
47MiniMax M3MiniMax1.4$0.23$0.96
48Qwen3.7 PlusAlibaba1.1$0.32$1.28
49Qwen3.6 PlusAlibaba0.9$0.33$1.95
50GLM 5.1Zhipu0.8$1.05$3.50
150 of 347
Page 1 of 7
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the Omniscience benchmark?

AA-Omniscience (Artificial Analysis) measures factual-knowledge reliability and calibration across 6,000 questions in six domains. Its Omniscience Index rewards correct answers and appropriate abstentions while penalising confident wrong answers, so a model that guesses badly can score below zero.

Which AI model scores highest on Omniscience?

Claude Fable 5 (Anthropic) leads with 43.3, from $10.00 per 1M input tokens.

How many models are ranked on Omniscience?

347 models carry a Omniscience score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI