Context
203K
Max output
131K
Serving providers
5
Cheapest input
$0.60 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: ARC Prize, Artificial Analysis and Hugging Face leaderboardsSpeed & latency
Output speed
62 t/s
tokens / second
Time to first token
1.42s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AIME 2026 | 96% | default | — | |||||||||||||
| APEX Agents | 14% | default | — | |||||||||||||
| ARC-AGI-2 | 5% | default | $0.270 | |||||||||||||
| CritPt | 2% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 621.5 | default | — | |||||||||||||
| ||||||||||||||||
| GPQA Diamond | 82% | default | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 29% | default | — | |||||||||||||
| ||||||||||||||||
| IFBench | 72% | default | — | |||||||||||||
| ||||||||||||||||
| Long-Context Reasoning | 71% | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience | 0.3 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.3 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.6 | default | — | |||||||||||||
| ||||||||||||||||
| SciCode | 46% | default | — | |||||||||||||
| ||||||||||||||||
| SWE-bench Verified | 78% | default | — | |||||||||||||
| TerminalBench Hard | 43% | default | — | |||||||||||||
| ||||||||||||||||
| τ²-bench | 98% | default | — | |||||||||||||
| ||||||||||||||||
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expandInput pricing runs from $0.60 to $1.00 per 1M tokens across 5 providers, so the dearest route costs 67% more than the cheapest for the same model.
| Serving provider | Input /1M | Output /1M | Endpoints | ||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepInfraCheapest | $0.60 | $2.08 | 1 | ||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| OpenRouterCheapest | $0.60 | $1.92 | 1 | ||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| Baseten | $0.95 | $3.15 | 1 | ||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| Zai | $1.00 | $3.20 | 1 | ||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| Novita | $1.00 | $3.20 | 1 | ||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Input down 25% since first tracked
Cost calculator
estimate your monthly spend on this model$286
estimated / month
Model IDs
copy the exact identifier for your platformdeepinfra/zai-org/GLM-5z-ai/glm-5baseten/zai-org/GLM-5zai/glm-5novita/zai-org/glm-5
Frequently asked questions
GLM 5 pricing, context and availabilityHow much does GLM 5 cost?
GLM 5 costs $0.60 per 1M input tokens and $1.92 per 1M output tokens at its cheapest provider via DeepInfra. Across 5 serving providers, input prices range from $0.60 to $1.00 per 1M tokens.
What is the context window of GLM 5?
GLM 5 accepts up to 203K tokens of context and can return up to 131K output tokens.
Which providers serve GLM 5?
GLM 5 is available from 5 serving providers, each with its own pricing and model ID. DeepInfra is currently the cheapest.
What can GLM 5 do?
GLM 5 supports tool use, reasoning, prompt caching and structured output.
Other Zhipu models
compare pricing across the Zhipu lineup Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI