Context
262K
Max output
262K
Serving providers
6
Cheapest input
$0.09 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.

Benchmarks

independent evaluations · source: ARC Prize and Artificial Analysis
Speed & latency
Output speed
62 t/s
tokens / second
Time to first token
2.36s
latency
Headline indices
22.1
of 100
3.8
of 100
Individual benchmarks · click a row for effort variants
BenchmarkBest scoreEffortCost / task
ARC-AGI-21%default$0.004
CritPt0%default
EvalComputeProxy292.7reasoning: true
GDPval544.4reasoning: true
GPQA Diamond79%reasoning: true
Humanity's Last Exam16%reasoning: true
IFBench51%reasoning: true
LiveCodeBench79%reasoning: true
Long-Context Reasoning71%reasoning: true
MMLU-Pro84%reasoning: true
Omniscience-44.0default
Omniscience Accuracy0.2reasoning: true
Omniscience Non Hallucination0.2default
SciCode42%reasoning: true
TerminalBench Hard15%default
TerminalBench v2.112%reasoning: true
τ-bench Banking8%reasoning: true
τ²-bench53%reasoning: true

Pricing by serving provider

standard tier · per 1M tokens · click a provider to expand

Input pricing runs from $0.09 to $3.00 per 1M tokens across 6 providers, so the dearest route costs 3233% more than the cheapest for the same model.

Serving providerInput /1MOutput /1MEndpoints
DeepInfraCheapest$0.09$0.551
EndpointStd inStd outCached inBatch in / outModel ID
StandardCheapest$0.09$0.55deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507
NovitaCheapest$0.09$0.581
Fireworks AI$0.22$0.881
Replicate$0.26$1.061
Scaleway$0.75$2.251
Crusoe$3.00$3.001

Price history

input + output $/1M since we started tracking
Input Output
Input down 31% since first tracked
$0.000$0.200$0.400$0.600Aug 22Oct 13May 1Aug 28$0.550$0.090

Cost calculator

estimate your monthly spend on this model
$62
estimated / month

Model IDs

copy the exact identifier for your platform
deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507novita/qwen/qwen3-235b-a22b-instruct-2507fireworks_ai/accounts/fireworks/models/qwen3-235b-a22b-instruct-2507replicate/qwen/qwen3-235b-a22b-instruct-2507scaleway/qwen/qwen3-235b-a22b-instruct-2507crusoe/Qwen/Qwen3-235B-A22B-Instruct-2507

Frequently asked questions

Qwen3 235B A22b Instruct 2507 pricing, context and availability

How much does Qwen3 235B A22b Instruct 2507 cost?

Qwen3 235B A22b Instruct 2507 costs $0.09 per 1M input tokens and $0.55 per 1M output tokens at its cheapest provider via DeepInfra. Across 6 serving providers, input prices range from $0.09 to $3.00 per 1M tokens.

What is the context window of Qwen3 235B A22b Instruct 2507?

Qwen3 235B A22b Instruct 2507 accepts up to 262K tokens of context and can return up to 262K output tokens.

Which providers serve Qwen3 235B A22b Instruct 2507?

Qwen3 235B A22b Instruct 2507 is available from 6 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.

What can Qwen3 235B A22b Instruct 2507 do?

Qwen3 235B A22b Instruct 2507 supports tool use and structured output.

Other Alibaba models

compare pricing across the Alibaba lineup
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI