Context
131K
Max output
131K
Serving providers
19
Cheapest input
$0.03 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis and Hugging Face leaderboards
Speed & latency
Output speed
158 t/s
tokens / second
Time to first token
0.85s
latency
Headline indices
30.4
of 100
13.4
of 100
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
AA-Briefcase0%
APEX Agents3%default
CritPt1%default
EvalComputeProxy164.5default
GDPval800.2default
GPQA Diamond78%default
Harvey Lab0.1default
Humanity's Last Exam20%default
IFBench69%default
IT-Bench SRE6%default
LiveCodeBench88%default
Long-Context Reasoning51%default
Mlcr Overall0.0default
MMLU-Pro81%default
Omniscience-49.3default
Omniscience Accuracy0.2default
Omniscience Non Hallucination0.1default
SciCode39%default
SWE-bench Pro16%default
SWE-bench Verified62%default
TerminalBench Hard23%default
TerminalBench v2.126%default
τ-bench Banking13%default
τ²-bench66%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.03 to $0.80 per 1M tokens across 19 providers, so the dearest route costs 2567% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
WandbCheapest$0.03$0.171
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$0.03$0.17wandb/openai/gpt-oss-120b
DeepInfra$0.04$0.171
OpenRouter$0.04$0.171
Novita$0.05$0.251
Ovhcloud$0.08$0.401
Baseten$0.10$0.501
Fireworks AI$0.15$0.601
Azure$0.15$0.601
Crusoe is dearest at $0.80 / $0.80

Price history

input + output $/1M since we started tracking
Input Output
Input down 80% since first tracked
$0.000$0.200$0.400$0.600Aug 5Oct 17Jun 10Aug 28$0.170$0.030

Cost calculator

estimate your monthly spend on this model
$20
estimated / month

Model IDs

copy the exact identifier for your platform
wandb/openai/gpt-oss-120bdeepinfra/openai/gpt-oss-120bopenai/gpt-oss-120bnovita/openai/gpt-oss-120bovhcloud/gpt-oss-120bbaseten/openai/gpt-oss-120bfireworks_ai/accounts/fireworks/models/gpt-oss-120bazure_ai/gpt-oss-120bgroq/openai/gpt-oss-120btogether_ai/openai/gpt-oss-120bwatsonx/openai/gpt-oss-120bscaleway/openai/gpt-oss-120btensormesh/openai/gpt-oss-120bdatabricks/databricks-gpt-oss-120breplicate/openai/gpt-oss-120bsambanova/gpt-oss-120bcloudflare/@cf/openai/gpt-oss-120bcerebras/gpt-oss-120bcrusoe/openai/gpt-oss-120b

Frequently asked questions

gpt-oss-120b pricing, context and availability

How much does gpt-oss-120b cost?

gpt-oss-120b costs $0.03 per 1M input tokens and $0.17 per 1M output tokens at its cheapest provider via Wandb. Across 19 serving providers, input prices range from $0.03 to $0.80 per 1M tokens.

What is the context window of gpt-oss-120b?

gpt-oss-120b accepts up to 131K tokens of context and can return up to 131K output tokens.

Which providers serve gpt-oss-120b?

gpt-oss-120b is available from 19 serving providers, each with its own pricing and model ID. Wandb is currently the cheapest.

What can gpt-oss-120b do?

gpt-oss-120b supports image input (vision), tool use, reasoning, prompt caching, structured output and web search.

Other OpenAI models

compare pricing across the OpenAI lineup

Popular comparisons

head-to-head pages featuring gpt-oss-120b
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI