Context
262K
Max output
66K
Serving providers
4
Cheapest input
$0.14 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: Artificial Analysis
Speed & latency
Output speed
146 t/s
tokens / second
Time to first token
2.14s
latency
Headline indices
37.0
of 100
11.8
of 100
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
CritPt1%default
EvalComputeProxy117.6reasoning: false
GDPval796.2reasoning: false
GPQA Diamond85%default
Humanity's Last Exam21%default
IFBench73%default
IT-Bench SRE22%default
Long-Context Reasoning68%default
MMMU-Pro73%default
Omniscience-48.1default
Omniscience Accuracy0.2default
Omniscience Non Hallucination0.1default
SciCode38%default
TerminalBench Hard27%default
TerminalBench v2.141%reasoning: false
τ-bench Banking5%reasoning: false
τ²-bench89%default

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.14 to $0.25 per 1M tokens across 4 providers, so the dearest route costs 79% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
DeepInfraCheapest$0.14$1.001
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$0.14$1.00$0.05deepinfra/Qwen/Qwen3.5-35B-A3B
Novita$0.25$2.001
Wandb$0.25$1.251
OpenRouter$0.25$1.251

Price history

input + output $/1M since we started tracking
Input Output
Input down 44% since first tracked
$0.000$0.500$1.00$1.50$2.00Mar 9Jul 29Aug 13Aug 28$1.00$0.140

Cost calculator

estimate your monthly spend on this model
$108
estimated / month

Model IDs

copy the exact identifier for your platform
deepinfra/Qwen/Qwen3.5-35B-A3Bnovita/qwen/qwen3.5-35b-a3bwandb/Qwen/Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b

Frequently asked questions

Qwen3.5-35B-A3B pricing, context and availability

How much does Qwen3.5-35B-A3B cost?

Qwen3.5-35B-A3B costs $0.14 per 1M input tokens and $1.00 per 1M output tokens at its cheapest provider via DeepInfra. Across 4 serving providers, input prices range from $0.14 to $0.25 per 1M tokens.

What is the context window of Qwen3.5-35B-A3B?

Qwen3.5-35B-A3B accepts up to 262K tokens of context and can return up to 66K output tokens.

Which providers serve Qwen3.5-35B-A3B?

Qwen3.5-35B-A3B is available from 4 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.

What can Qwen3.5-35B-A3B do?

Qwen3.5-35B-A3B supports image input (vision), tool use, reasoning, prompt caching and structured output.

Other Alibaba models

compare pricing across the Alibaba lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI