Context
256K
Max output
256K
Serving providers
1
Cheapest input
$0.50 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial Analysis and Hugging Face leaderboardsSpeed & latency
Output speed
143 t/s
tokens / second
Time to first token
1.80s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |
|---|---|---|---|---|
| AIME 2026 | 90% | default | — | |
| APEX Agents | 2% | default | — | |
| CritPt | 3% | default | — | |
| EvalComputeProxy | 1042.6 | default | — | |
| GDPval | 698.1 | default | — | |
| GPQA Diamond | 80% | default | — | |
| Humanity's Last Exam | 21% | default | — | |
| IFBench | 71% | default | — | |
| IT-Bench SRE | 1% | default | — | |
| Long-Context Reasoning | 60% | default | — | |
| Mlcr Overall | 0.0 | default | — | |
| Omniscience | -41.5 | default | — | |
| Omniscience Accuracy | 0.2 | default | — | |
| Omniscience Non Hallucination | 0.1 | default | — | |
| SciCode | 36% | default | — | |
| SWE-bench Verified | 60% | default | — | |
| TerminalBench Hard | 29% | default | — | |
| TerminalBench v2.1 | 39% | default | — | |
| τ-bench Banking | 10% | default | — | |
| τ²-bench | 68% | default | — |
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expand| Serving provider | Input /1M | Output /1M | Endpoints | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CloudflareCheapest | $0.50 | $1.50 | 1 | ||||||||||||||||
| |||||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Price history is accruing. We record a point each day a price changes; the full backfill lands shortly.
Cost calculator
estimate your monthly spend on this model$220
estimated / month
Model IDs
copy the exact identifier for your platformcloudflare/@cf/nvidia/nemotron-3-120b-a12b
Frequently asked questions
Nemotron 3 120B A12b pricing, context and availabilityHow much does Nemotron 3 120B A12b cost?
Nemotron 3 120B A12b costs $0.50 per 1M input tokens and $1.50 per 1M output tokens at its cheapest provider via Cloudflare.
What is the context window of Nemotron 3 120B A12b?
Nemotron 3 120B A12b accepts up to 256K tokens of context and can return up to 256K output tokens.
What can Nemotron 3 120B A12b do?
Nemotron 3 120B A12b supports tool use and reasoning.
Other NVIDIA models
compare pricing across the NVIDIA lineupNvidia Nemotron Nano 9B$0.04/1M · 3 providersNemotron 3 Nano 30B A3B$0.05/1M · 3 providersNvidia Nemotron Nano 12B$0.20/1M · 2 providersNvidia Nemotron 3 Ultra 550B A55b$0.50/1M · 2 providersNemotron 3 Ultra$0.50/1M · 2 providersNemotron 3 Nano 30B A3B (free)$0.00/1M · 1 providerNemotron 3 Nano Omni (free)$0.00/1M · 1 providerNemotron 3 Super (free)$0.00/1M · 1 provider
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI