Context
41K
Max output
41K
Serving providers
4
Cheapest input
$0.08 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial AnalysisSpeed & latency
Output speed
60 t/s
tokens / second
Time to first token
2.69s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CritPt | 0% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 96.7 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| GDPval | 235.8 | reasoning: true | — | |||||||||||||
| GPQA Diamond | 60% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 5% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| IFBench | 41% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| LiveCodeBench | 52% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Long-Context Reasoning | 0% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| MMLU-Pro | 77% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience | -49.0 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.2 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.2 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| SciCode | 32% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| TerminalBench Hard | 5% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench v2.1 | 5% | reasoning: true | — | |||||||||||||
| τ-bench Banking | 6% | reasoning: true | — | |||||||||||||
| τ²-bench | 35% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
Pricing detail
cost beyond the standard rate · source: models.devReasoning tokens
$4.20
per 1M · thinking output
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expandInput pricing runs from $0.08 to $0.20 per 1M tokens across 4 providers, so the dearest route costs 150% more than the cheapest for the same model.
| Serving provider | Input /1M | Output /1M | Endpoints | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| NebiusCheapest | $0.08 | $0.24 | 1 | ||||||||||||||||
| |||||||||||||||||||
| DeepInfra | $0.12 | $0.24 | 1 | ||||||||||||||||
| |||||||||||||||||||
| OpenRouter | $0.12 | $0.24 | 1 | ||||||||||||||||
| |||||||||||||||||||
| Fireworks AI | $0.20 | $0.20 | 1 | ||||||||||||||||
| |||||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Cost calculator
estimate your monthly spend on this model$35
estimated / month
Model IDs
copy the exact identifier for your platformnebius/Qwen/Qwen3-14Bdeepinfra/Qwen/Qwen3-14Bqwen/qwen3-14bfireworks_ai/accounts/fireworks/models/qwen3-14b
Frequently asked questions
Qwen3 14B pricing, context and availabilityHow much does Qwen3 14B cost?
Qwen3 14B costs $0.08 per 1M input tokens and $0.20 per 1M output tokens at its cheapest provider via Nebius. Across 4 serving providers, input prices range from $0.08 to $0.20 per 1M tokens.
What is the context window of Qwen3 14B?
Qwen3 14B accepts up to 41K tokens of context and can return up to 41K output tokens.
Which providers serve Qwen3 14B?
Qwen3 14B is available from 4 serving providers, each with its own pricing and model ID. Nebius is currently the cheapest.
What can Qwen3 14B do?
Qwen3 14B supports tool use.
Other Alibaba models
compare pricing across the Alibaba lineupQwen3 32B$0.08/1M · 8 providersQwen3.6 35B A3B$0.10/1M · 7 providersQwQ 32B$0.15/1M · 7 providersQwen3 Next 80B A3B Instruct$0.09/1M · 6 providersQwen3 235B A22b Instruct 2507$0.09/1M · 6 providersQwen3 Next 80B A3B Thinking$0.14/1M · 6 providersQwen3.6 27B$0.15/1M · 6 providersQwen3 30B A3B$0.09/1M · 5 providers
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI