Context
131K
Max output
131K
Serving providers
13
Cheapest input
$0.01 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial Analysis and Hugging Face leaderboardsSpeed & latency
Output speed
124 t/s
tokens / second
Time to first token
1.05s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AA-Briefcase | 0% | — | — | |||||||||||||
| APEX Agents | 1% | default | — | |||||||||||||
| CritPt | 1% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 200.1 | default | — | |||||||||||||
| ||||||||||||||||
| GDPval | 566.6 | default | — | |||||||||||||
| GPQA Diamond | 69% | default | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 11% | default | — | |||||||||||||
| ||||||||||||||||
| IFBench | 65% | default | — | |||||||||||||
| ||||||||||||||||
| LiveCodeBench | 78% | default | — | |||||||||||||
| ||||||||||||||||
| Long-Context Reasoning | 33% | default | — | |||||||||||||
| ||||||||||||||||
| Mlcr Overall | 0.0 | default | — | |||||||||||||
| MMLU-Pro | 75% | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience | -58.5 | low | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.2 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.1 | low | — | |||||||||||||
| ||||||||||||||||
| SciCode | 34% | default | — | |||||||||||||
| ||||||||||||||||
| SWE-bench Verified | 61% | default | — | |||||||||||||
| TerminalBench Hard | 11% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench v2.1 | 14% | default | — | |||||||||||||
| τ-bench Banking | 7% | default | — | |||||||||||||
| τ²-bench | 60% | default | — | |||||||||||||
| ||||||||||||||||
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expandInput pricing runs from $0.01 to $0.20 per 1M tokens across 13 providers, so the dearest route costs 1279% more than the cheapest for the same model.
| Serving provider | Input /1M | Output /1M | Endpoints | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DarkbloomCheapest | $0.01 | $0.07 | 1 | ||||||||||||||||
| |||||||||||||||||||
| DeepInfra | $0.03 | $0.14 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Wandb | $0.03 | $0.13 | 1 | ||||||||||||||||
| |||||||||||||||||||
| OpenRouter | $0.03 | $0.13 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Ovhcloud | $0.04 | $0.15 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Novita | $0.04 | $0.15 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Together AI | $0.05 | $0.20 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Databricks | $0.07 | $0.30 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Cloudflare is dearest at $0.20 / $0.30 | |||||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Input down 71% since first tracked
Cost calculator
estimate your monthly spend on this model$9
estimated / month
Model IDs
copy the exact identifier for your platformdarkbloom/gpt-oss-20bdeepinfra/openai/gpt-oss-20bwandb/openai/gpt-oss-20bopenai/gpt-oss-20bovhcloud/gpt-oss-20bnovita/openai/gpt-oss-20btogether_ai/openai/gpt-oss-20bdatabricks/databricks-gpt-oss-20bfireworks_ai/accounts/fireworks/models/gpt-oss-20btensormesh/openai/gpt-oss-20bgroq/openai/gpt-oss-20breplicate/openai/gpt-oss-20bcloudflare/@cf/openai/gpt-oss-20b
Frequently asked questions
gpt-oss-20b pricing, context and availabilityHow much does gpt-oss-20b cost?
gpt-oss-20b costs $0.01 per 1M input tokens and $0.07 per 1M output tokens at its cheapest provider via Darkbloom. Across 13 serving providers, input prices range from $0.01 to $0.20 per 1M tokens.
What is the context window of gpt-oss-20b?
gpt-oss-20b accepts up to 131K tokens of context and can return up to 131K output tokens.
Which providers serve gpt-oss-20b?
gpt-oss-20b is available from 13 serving providers, each with its own pricing and model ID. Darkbloom is currently the cheapest.
What can gpt-oss-20b do?
gpt-oss-20b supports image input (vision), tool use, reasoning, prompt caching, structured output and web search.
Other OpenAI models
compare pricing across the OpenAI lineupPopular comparisons
head-to-head pages featuring gpt-oss-20b Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI