This model was retired on 2026-03-31. The pricing below is the last-known rate, kept for migration reference.

Llama 4 Maverick 17B 128e Instruct FP8Deprecated

MetaReleasedApr 5, 2025KnowledgeAug 2024
VisionTool useStructured output
Context
1M
Max output
1M
Serving providers
6
Cheapest input
$0.05 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: ARC Prize
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
ARC-AGI-20%default$0.012

Pricing by serving provider

last-known · per 1M tokens · click a provider to expand
Serving providerInput /1MOutput /1MEndpoints
Lambda AI$0.05$0.101
DeepInfra$0.20$0.801
Together AI$0.27$0.851
Novita$0.27$0.851
OCI$0.72$0.721
Azure$1.41$0.351

Price history

input + output $/1M since we started tracking
Input Output
Input down 96% since first tracked
$0.000$0.500$1.00$1.50May 9Jul 21Jan 13Aug 28$0.100$0.050

Cost calculator

estimate your monthly spend on this model
$18
estimated / month

Model IDs

copy the exact identifier for your platform
lambda_ai/llama-4-maverick-17b-128e-instruct-fp8deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8together_ai/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8novita/meta-llama/llama-4-maverick-17b-128e-instruct-fp8oci/meta.llama-4-maverick-17b-128e-instruct-fp8azure_ai/Llama-4-Maverick-17B-128E-Instruct-FP8

Frequently asked questions

Llama 4 Maverick 17B 128e Instruct FP8 pricing, context and availability

Is Llama 4 Maverick 17B 128e Instruct FP8 still available?

Llama 4 Maverick 17B 128e Instruct FP8 was retired on 2026-03-31. The pricing on this page is the last-known rate, kept for migration reference.

How much did Llama 4 Maverick 17B 128e Instruct FP8 cost?

Llama 4 Maverick 17B 128e Instruct FP8's last-known pricing, before it was retired on 2026-03-31, was $0.05 per 1M input tokens and $0.10 per 1M output tokens.

What was the context window of Llama 4 Maverick 17B 128e Instruct FP8?

Llama 4 Maverick 17B 128e Instruct FP8 had a 1M token context window and could return up to 1M output tokens.

What could Llama 4 Maverick 17B 128e Instruct FP8 do?

Llama 4 Maverick 17B 128e Instruct FP8 supported image input (vision), tool use and structured output.

Other Meta models

compare pricing across the Meta lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI