Context
262K
Max output
66K
Serving providers
4
Cheapest input
$0.14 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial AnalysisSpeed & latency
Output speed
146 t/s
tokens / second
Time to first token
2.14s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CritPt | 1% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 117.6 | reasoning: false | — | |||||||||||||
| ||||||||||||||||
| GDPval | 796.2 | reasoning: false | — | |||||||||||||
| GPQA Diamond | 85% | default | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 21% | default | — | |||||||||||||
| ||||||||||||||||
| IFBench | 73% | default | — | |||||||||||||
| ||||||||||||||||
| IT-Bench SRE | 22% | default | — | |||||||||||||
| Long-Context Reasoning | 68% | default | — | |||||||||||||
| ||||||||||||||||
| MMMU-Pro | 73% | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience | -48.1 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.2 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.1 | default | — | |||||||||||||
| ||||||||||||||||
| SciCode | 38% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench Hard | 27% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench v2.1 | 41% | reasoning: false | — | |||||||||||||
| τ-bench Banking | 5% | reasoning: false | — | |||||||||||||
| τ²-bench | 89% | default | — | |||||||||||||
| ||||||||||||||||
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expandInput pricing runs from $0.14 to $0.25 per 1M tokens across 4 providers, so the dearest route costs 79% more than the cheapest for the same model.
| Serving provider | Input /1M | Output /1M | Endpoints | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepInfraCheapest | $0.14 | $1.00 | 1 | |||||||||||||
| ||||||||||||||||
| Novita | $0.25 | $2.00 | 1 | |||||||||||||
| ||||||||||||||||
| Wandb | $0.25 | $1.25 | 1 | |||||||||||||
| ||||||||||||||||
| OpenRouter | $0.25 | $1.25 | 1 | |||||||||||||
| ||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Input down 44% since first tracked
Cost calculator
estimate your monthly spend on this model$108
estimated / month
Model IDs
copy the exact identifier for your platformdeepinfra/Qwen/Qwen3.5-35B-A3Bnovita/qwen/qwen3.5-35b-a3bwandb/Qwen/Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b
Frequently asked questions
Qwen3.5-35B-A3B pricing, context and availabilityHow much does Qwen3.5-35B-A3B cost?
Qwen3.5-35B-A3B costs $0.14 per 1M input tokens and $1.00 per 1M output tokens at its cheapest provider via DeepInfra. Across 4 serving providers, input prices range from $0.14 to $0.25 per 1M tokens.
What is the context window of Qwen3.5-35B-A3B?
Qwen3.5-35B-A3B accepts up to 262K tokens of context and can return up to 66K output tokens.
Which providers serve Qwen3.5-35B-A3B?
Qwen3.5-35B-A3B is available from 4 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.
What can Qwen3.5-35B-A3B do?
Qwen3.5-35B-A3B supports image input (vision), tool use, reasoning, prompt caching and structured output.
Other Alibaba models
compare pricing across the Alibaba lineupQwen3 32B$0.08/1M · 8 providersQwen3.6 35B A3B$0.10/1M · 7 providersQwQ 32B$0.15/1M · 7 providersQwen3 Next 80B A3B Instruct$0.09/1M · 6 providersQwen3 235B A22b Instruct 2507$0.09/1M · 6 providersQwen3 Next 80B A3B Thinking$0.14/1M · 6 providersQwen3.6 27B$0.15/1M · 6 providersQwen3 30B A3B$0.09/1M · 5 providers
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI