Context
262K
Max output
n/a
Serving providers
2
Cheapest input
$0.50 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis and Hugging Face leaderboards
Speed & latency
Output speed
105 t/s
tokens / second
Time to first token
3.57s
latency
Headline indices
49.3
of 100
27.5
of 100
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
Aa Analyst Agent0.1default
Automation Bench0.1default
CritPt3%default
EvalComputeProxy2065.9default
GDPval1163.0default
GPQA Diamond87%default
Harvey Lab0.8default
Humanity's Last Exam28%default
IFBench81%default
Long-Context Reasoning71%default
Mlcr Overall0.1default
Omniscience-0.4default
Omniscience Accuracy0.2default
Omniscience Non Hallucination0.7default
SciCode40%default
SWE-bench Verified72%default
TerminalBench Hard36%default
TerminalBench v2.154%default
τ-bench Banking14%default
τ²-bench83%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.50 to $0.75 per 1M tokens across 2 providers, so the dearest route costs 50% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
DeepInfraCheapest$0.50$2.201
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$0.50$2.20$0.10deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
Wandb$0.75$2.751

Price history

input + output $/1M since we started tracking
Input Output

Price history is accruing. We record a point each day a price changes; the full backfill lands shortly.

Cost calculator

estimate your monthly spend on this model
$276
estimated / month

Model IDs

copy the exact identifier for your platform
deepinfra/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55Bwandb/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B

Frequently asked questions

Nvidia Nemotron 3 Ultra 550B A55b pricing, context and availability

How much does Nvidia Nemotron 3 Ultra 550B A55b cost?

Nvidia Nemotron 3 Ultra 550B A55b costs $0.50 per 1M input tokens and $2.20 per 1M output tokens at its cheapest provider via DeepInfra. Across 2 serving providers, input prices range from $0.50 to $0.75 per 1M tokens.

What is the context window of Nvidia Nemotron 3 Ultra 550B A55b?

Nvidia Nemotron 3 Ultra 550B A55b accepts up to 262K tokens of context.

Which providers serve Nvidia Nemotron 3 Ultra 550B A55b?

Nvidia Nemotron 3 Ultra 550B A55b is available from 2 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.

What can Nvidia Nemotron 3 Ultra 550B A55b do?

Nvidia Nemotron 3 Ultra 550B A55b supports image input (vision), tool use, reasoning, prompt caching and structured output.

Other NVIDIA models

compare pricing across the NVIDIA lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI