Context
131K
Max output
131K
Serving providers
4
Cheapest input
$0.40 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis
Speed & latency
Output speed
55 t/s
tokens / second
Time to first token
1.33s
latency
Headline indices
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
CritPt0%default
EvalComputeProxy15.0default
GPQA Diamond41%default
Humanity's Last Exam4%default
IFBench34%default
LiveCodeBench23%default
Long-Context Reasoning8%default
MMLU-Pro68%default
Omniscience-43.1default
Omniscience Accuracy0.2default
Omniscience Non Hallucination0.2default
SciCode27%default
TerminalBench Hard3%default
τ²-bench15%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.40 to $1.00 per 1M tokens across 4 providers, so the dearest route costs 150% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
OpenRouterCheapest$0.40$0.401
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$0.40$0.40meta-llama/llama-3.1-70b-instruct
OCI$0.72$0.721
Wandb$0.80$0.801
Perplexity$1.00$1.001

Price history

input + output $/1M since we started tracking
Input Output
Input down 60% since first tracked
$0.000$0.500$1.00Aug 14Apr 4Jun 22Aug 28$0.400$0.400

Cost calculator

estimate your monthly spend on this model
$112
estimated / month

Model IDs

copy the exact identifier for your platform
meta-llama/llama-3.1-70b-instructoci/meta.llama-3.1-70b-instructwandb/meta-llama/Llama-3.1-70B-Instructperplexity/llama-3.1-70b-instruct

Frequently asked questions

Llama 3.1 70B Instruct pricing, context and availability

How much does Llama 3.1 70B Instruct cost?

Llama 3.1 70B Instruct costs $0.40 per 1M input tokens and $0.40 per 1M output tokens at its cheapest provider via OpenRouter. Across 4 serving providers, input prices range from $0.40 to $1.00 per 1M tokens.

What is the context window of Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct accepts up to 131K tokens of context and can return up to 131K output tokens.

Which providers serve Llama 3.1 70B Instruct?

Llama 3.1 70B Instruct is available from 4 serving providers, each with its own pricing and model ID. OpenRouter is currently the cheapest.

What can Llama 3.1 70B Instruct do?

Llama 3.1 70B Instruct supports tool use.

Other Meta models

compare pricing across the Meta lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI