Context
131K
Max output
131K
Serving providers
4
Cheapest input
$0.40 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial AnalysisSpeed & latency
Output speed
55 t/s
tokens / second
Time to first token
1.33s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |
|---|---|---|---|---|
| CritPt | 0% | default | — | |
| EvalComputeProxy | 15.0 | default | — | |
| GPQA Diamond | 41% | default | — | |
| Humanity's Last Exam | 4% | default | — | |
| IFBench | 34% | default | — | |
| LiveCodeBench | 23% | default | — | |
| Long-Context Reasoning | 8% | default | — | |
| MMLU-Pro | 68% | default | — | |
| Omniscience | -43.1 | default | — | |
| Omniscience Accuracy | 0.2 | default | — | |
| Omniscience Non Hallucination | 0.2 | default | — | |
| SciCode | 27% | default | — | |
| TerminalBench Hard | 3% | default | — | |
| τ²-bench | 15% | default | — |
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expandInput pricing runs from $0.40 to $1.00 per 1M tokens across 4 providers, so the dearest route costs 150% more than the cheapest for the same model.
| Serving provider | Input /1M | Output /1M | Endpoints | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenRouterCheapest | $0.40 | $0.40 | 1 | ||||||||||||||||
| |||||||||||||||||||
| OCI | $0.72 | $0.72 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Wandb | $0.80 | $0.80 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Perplexity | $1.00 | $1.00 | 1 | ||||||||||||||||
| |||||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Input down 60% since first tracked
Cost calculator
estimate your monthly spend on this model$112
estimated / month
Model IDs
copy the exact identifier for your platformmeta-llama/llama-3.1-70b-instructoci/meta.llama-3.1-70b-instructwandb/meta-llama/Llama-3.1-70B-Instructperplexity/llama-3.1-70b-instruct
Frequently asked questions
Llama 3.1 70B Instruct pricing, context and availabilityHow much does Llama 3.1 70B Instruct cost?
Llama 3.1 70B Instruct costs $0.40 per 1M input tokens and $0.40 per 1M output tokens at its cheapest provider via OpenRouter. Across 4 serving providers, input prices range from $0.40 to $1.00 per 1M tokens.
What is the context window of Llama 3.1 70B Instruct?
Llama 3.1 70B Instruct accepts up to 131K tokens of context and can return up to 131K output tokens.
Which providers serve Llama 3.1 70B Instruct?
Llama 3.1 70B Instruct is available from 4 serving providers, each with its own pricing and model ID. OpenRouter is currently the cheapest.
What can Llama 3.1 70B Instruct do?
Llama 3.1 70B Instruct supports tool use.
Other Meta models
compare pricing across the Meta lineupLlama 3.3 70B Instruct$0.12/1M · 13 providersLlama 4 Scout 17B 16e Instruct$0.05/1M · 10 providersLlama 3.1 8B Instruct$0.02/1M · 8 providersMeta Llama 3.1 8B Instruct$0.02/1M · 6 providersLlama 3.2 3B Instruct$0.02/1M · 6 providersLlama 4 Maverick 17B 128e Instruct FP8$0.05/1M · 6 providersLlama 3.2 11B Vision Instruct$0.05/1M · 5 providersMeta Llama 3.1 70B Instruct$0.12/1M · 5 providers
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI