Context
1M
Max output
1M
Serving providers
3
Cheapest input
$2.00 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis and Hugging Face leaderboards
Speed & latency
Output speed
27 t/s
tokens / second
Time to first token
2.50s
latency
Headline indices
71.9
of 100
57.1
of 100
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
CritPt20%default
EvalComputeProxy26108.9default
GDPval1720.4default
GPQA Diamond94%default
Humanity's Last Exam42%default
Long-Context Reasoning75%default
Mlcr Overall0default
Omniscience4.3default
Omniscience Accuracy0.3default
Omniscience Non Hallucination0.6default
SciCode52%default
SWE-bench Pro68%default
TerminalBench v2.182%default
τ-bench Banking49%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $2.00 to $2.50 per 1M tokens across 3 providers, so the dearest route costs 25% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
DeepInfraCheapest$2.00$6.001
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$2.00$6.00$0.20deepinfra/Qwen/Qwen3.8-2.4T-A95B
OpenRouterCheapest$2.00$6.001
Together AI$2.50$6.251

Price history

input + output $/1M since we started tracking
Input Output
$0.000$2.00$4.00$6.00Aug 13Aug 14Aug 26Aug 28$6.00$2.00

Cost calculator

estimate your monthly spend on this model
$880
estimated / month

Model IDs

copy the exact identifier for your platform
deepinfra/Qwen/Qwen3.8-2.4T-A95Bqwen/qwen3.8-2.4t-a95btogether_ai/Qwen/Qwen3.8-2.4T-A95B

Frequently asked questions

Qwen3.8 2.4T A95B pricing, context and availability

How much does Qwen3.8 2.4T A95B cost?

Qwen3.8 2.4T A95B costs $2.00 per 1M input tokens and $6.00 per 1M output tokens at its cheapest provider via DeepInfra. Across 3 serving providers, input prices range from $2.00 to $2.50 per 1M tokens.

What is the context window of Qwen3.8 2.4T A95B?

Qwen3.8 2.4T A95B accepts up to 1M tokens of context and can return up to 1M output tokens.

Which providers serve Qwen3.8 2.4T A95B?

Qwen3.8 2.4T A95B is available from 3 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.

What can Qwen3.8 2.4T A95B do?

Qwen3.8 2.4T A95B supports tool use, reasoning, prompt caching and structured output.

Other Alibaba models

compare pricing across the Alibaba lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI