Context
131K
Max output
131K
Serving providers
1
Cheapest input
$0.10 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand
Serving providerInput /1MOutput /1MEndpoints
DeepInfraCheapest$0.10$0.401
ComponentUnitStandardBatchCached
Text input/1M tok$0.1000
Text output/1M tok$0.4000

Price history

input + output $/1M since we started tracking
Input Output
$0.000$0.100$0.200$0.300$0.400Sep 26Jun 22$0.400$0.100

Cost calculator

estimate your monthly spend on this model
$52
estimated / month

Model IDs

copy the exact identifier for your platform
deepinfra/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5

Frequently asked questions

Llama 3.3 Nemotron Super 49B V1 5 pricing, context and availability

How much does Llama 3.3 Nemotron Super 49B V1 5 cost?

Llama 3.3 Nemotron Super 49B V1 5 costs $0.10 per 1M input tokens and $0.40 per 1M output tokens at its cheapest provider via DeepInfra.

What is the context window of Llama 3.3 Nemotron Super 49B V1 5?

Llama 3.3 Nemotron Super 49B V1 5 accepts up to 131K tokens of context and can return up to 131K output tokens.

What can Llama 3.3 Nemotron Super 49B V1 5 do?

Llama 3.3 Nemotron Super 49B V1 5 supports tool use.

Other NVIDIA models

compare pricing across the NVIDIA lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI