Context
n/a
Max output
n/a
Serving providers
0
Cheapest input
n/a /1M
Benchmarks
independent evaluations · source: Artificial AnalysisHeadline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CritPt | 1% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 783.9 | default | — | |||||||||||||
| ||||||||||||||||
| GDPval | 592.5 | default | — | |||||||||||||
| GPQA Diamond | 78% | default | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 14% | default | — | |||||||||||||
| ||||||||||||||||
| IFBench | 65% | default | — | |||||||||||||
| ||||||||||||||||
| LiveCodeBench | 77% | default | — | |||||||||||||
| Long-Context Reasoning | 59% | default | — | |||||||||||||
| ||||||||||||||||
| MMLU-Pro | 84% | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience | -58.0 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.2 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.1 | default | — | |||||||||||||
| ||||||||||||||||
| SciCode | 36% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench Hard | 23% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench v2.1 | 30% | default | — | |||||||||||||
| τ-bench Banking | 14% | default | — | |||||||||||||
| τ²-bench | 74% | default | — | |||||||||||||
| ||||||||||||||||
Pricing by serving provider
availabilityNo public price yet. Not offered on-demand at any tracked provider.
Price history
input + output $/1M since we started tracking Input Output
Price history is accruing. We record a point each day a price changes; the full backfill lands shortly.
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI