Context
n/a
Max output
n/a
Serving providers
0
Cheapest input
n/a /1M
Benchmarks
independent evaluations · source: Artificial Analysis and Hugging Face leaderboardsHeadline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |
|---|---|---|---|---|
| AIME 2026 | 97% | default | — | |
| CritPt | 9% | default | — | |
| EvalComputeProxy | 4360.1 | default | — | |
| GDPval | 1109.4 | default | — | |
| GPQA Diamond | 86% | default | — | |
| Humanity's Last Exam | 30% | default | — | |
| Long-Context Reasoning | 66% | default | — | |
| Omniscience | -8.2 | default | — | |
| Omniscience Accuracy | 0.2 | default | — | |
| Omniscience Non Hallucination | 0.7 | default | — | |
| SciCode | 39% | default | — | |
| TerminalBench v2.1 | 39% | default | — | |
| τ-bench Banking | 16% | default | — |
Pricing by serving provider
availabilityNo public price yet. Not offered on-demand at any tracked provider.
Price history
input + output $/1M since we started tracking Input Output
Price history is accruing. We record a point each day a price changes; the full backfill lands shortly.
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI