This model was retired on 2027-01-26. The pricing below is the last-known rate, kept for migration reference.
Context
262K
Max output
262K
Serving providers
8
Cheapest input
$0.45 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: ARC Prize, Artificial Analysis and Hugging Face leaderboardsSpeed & latency
Output speed
66 t/s
tokens / second
Time to first token
2.80s
latency
Headline indices
Intelligence Index
36.0
of 100
Coding Index
46.8
of 100
Agentic Index
21.7
of 100
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AIME 2026 | 96% | default | — | |||||||||||||
| APEX Agents | 12% | default | — | |||||||||||||
| ARC-AGI-2 | 12% | default | $0.280 | |||||||||||||
| CritPt | 3% | default | — | |||||||||||||
| ||||||||||||||||
| EvalComputeProxy | 648.1 | default | — | |||||||||||||
| ||||||||||||||||
| GDPval | 1006.1 | default | — | |||||||||||||
| GPQA Diamond | 88% | default | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 31% | default | — | |||||||||||||
| ||||||||||||||||
| IFBench | 70% | default | — | |||||||||||||
| ||||||||||||||||
| Long-Context Reasoning | 73% | default | — | |||||||||||||
| ||||||||||||||||
| MMMU-Pro | 75% | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience | -7.3 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.4 | default | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.5 | reasoning: false | — | |||||||||||||
| ||||||||||||||||
| SciCode | 49% | default | — | |||||||||||||
| ||||||||||||||||
| SWE-bench Pro | 51% | default | — | |||||||||||||
| SWE-bench Verified | 71% | default | — | |||||||||||||
| TerminalBench Hard | 35% | default | — | |||||||||||||
| ||||||||||||||||
| TerminalBench v2.1 | 46% | default | — | |||||||||||||
| τ-bench Banking | 14% | default | — | |||||||||||||
| τ²-bench | 96% | default | — | |||||||||||||
| ||||||||||||||||
Pricing by serving provider
last-known · per 1M tokens · click a provider to expand| Serving provider | Input /1M | Output /1M | Endpoints | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepInfra | $0.45 | $2.25 | 1 | |||||||||||||
| ||||||||||||||||
| Together AI | $0.50 | $2.80 | 1 | |||||||||||||
| ||||||||||||||||
| Azure | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
| Baseten | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
| Moonshot | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
| Wandb | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
| Novita | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
| OpenRouter | $0.60 | $3.00 | 1 | |||||||||||||
| ||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Input down 25% since first tracked
Cost calculator
estimate your monthly spend on this model$270
estimated / month
Model IDs
copy the exact identifier for your platformdeepinfra/moonshotai/Kimi-K2.5together_ai/moonshotai/Kimi-K2.5azure_ai/kimi-k2.5baseten/moonshotai/Kimi-K2.5moonshot/kimi-k2.5wandb/moonshotai/Kimi-K2.5novita/moonshotai/kimi-k2.5moonshotai/kimi-k2.5
Frequently asked questions
Kimi K2.5 pricing, context and availabilityIs Kimi K2.5 still available?
Kimi K2.5 was retired on 2027-01-26. The pricing on this page is the last-known rate, kept for migration reference.
How much did Kimi K2.5 cost?
Kimi K2.5's last-known pricing, before it was retired on 2027-01-26, was $0.45 per 1M input tokens and $2.25 per 1M output tokens.
What was the context window of Kimi K2.5?
Kimi K2.5 had a 262K token context window and could return up to 262K output tokens.
What could Kimi K2.5 do?
Kimi K2.5 supported image input (vision), tool use, reasoning, prompt caching and structured output.
Other Moonshot models
compare pricing across the Moonshot lineup Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI