Context
131K
Max output
131K
Serving providers
13
Cheapest input
$0.12 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis
Speed & latency
Output speed
88 t/s
tokens / second
Time to first token
1.65s
latency
Headline indices
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
CritPt0%default
EvalComputeProxy1125.1default
GDPval96.7default
GPQA Diamond50%default
Humanity's Last Exam4%default
IFBench47%default
IT-Bench SRE1%default
LiveCodeBench29%default
Long-Context Reasoning16%default
MMLU-Pro71%default
Omniscience-54.2default
Omniscience Accuracy0.2default
Omniscience Non Hallucination0.1default
SciCode26%default
TerminalBench Hard3%default
TerminalBench v2.15%default
τ-bench Banking1%default
τ²-bench27%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.12 to $0.90 per 1M tokens across 13 providers, so the dearest route costs 650% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
HyperbolicCheapest$0.12$0.301
ComponentUnitStandardBatchCached
Text input/1M tok$0.1200
Text output/1M tok$0.3000
Nebius$0.13$0.401
Novita$0.14$0.401
Nscale$0.20$0.201
Crusoe$0.20$0.201
DeepInfra$0.23$0.401
Databricks$0.50$1.501
Azure$0.71$0.711
Scaleway is dearest at $0.90 / $0.90

Price history

input + output $/1M since we started tracking
Input Output
Input down 83% since first tracked
$0.000$0.200$0.400$0.600$0.800Dec 18Aug 7May 1Aug 28$0.200$0.120

Cost calculator

estimate your monthly spend on this model
$48
estimated / month

Model IDs

copy the exact identifier for your platform
hyperbolic/meta-llama/Llama-3.3-70B-Instructnebius/meta-llama/Llama-3.3-70B-Instructnovita/meta-llama/llama-3.3-70b-instructnscale/meta-llama/Llama-3.3-70B-Instructcrusoe/meta-llama/Llama-3.3-70B-Instructdeepinfra/meta-llama/Llama-3.3-70B-Instructdatabricks/databricks-meta-llama-3-3-70b-instructazure_ai/Llama-3.3-70B-Instructwandb/meta-llama/Llama-3.3-70B-Instructwatsonx/meta-llama/llama-3-3-70b-instructmeta-llama/llama-3.3-70b-instructoci/meta.llama-3.3-70b-instructscaleway/meta/llama-3.3-70b-instruct

Frequently asked questions

Llama 3.3 70B Instruct pricing, context and availability

How much does Llama 3.3 70B Instruct cost?

Llama 3.3 70B Instruct costs $0.12 per 1M input tokens and $0.20 per 1M output tokens at its cheapest provider via Hyperbolic. Across 13 serving providers, input prices range from $0.12 to $0.90 per 1M tokens.

What is the context window of Llama 3.3 70B Instruct?

Llama 3.3 70B Instruct accepts up to 131K tokens of context and can return up to 131K output tokens.

Which providers serve Llama 3.3 70B Instruct?

Llama 3.3 70B Instruct is available from 13 serving providers, each with its own pricing and model ID. Hyperbolic is currently the cheapest.

What can Llama 3.3 70B Instruct do?

Llama 3.3 70B Instruct supports tool use and structured output.

Other Meta models

compare pricing across the Meta lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI