τ-bench Banking benchmark
165 AI models ranked on τ-bench Banking. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.
Top model today
Qwen3.8 Max (Alibaba) leads at 51%, from $1.65 per 1M input.
Data via GitHub: sierra-research/tau2-bench →Models ranked
165
on this benchmark
Top score
51%
current leader
Type
Single eval
benchmark
Updated
29 Aug 2026
last source fetch
Cost vs τ-bench Banking
111 priced modelsEach labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.
Models ranked on τ-bench Banking
highest score first · 165 models| # | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 Max | 51% | $1.65 | $4.95 | ||||||||||||||||||||||||||||||
| 2 | Grok 4.6 | 51% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 3 | GLM 5.3 | 50% | $1.40 | $4.40 | ||||||||||||||||||||||||||||||
| 4 | Qwen3.8 2.4T A95B | 49% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 5 | Qwen3.8 27B | 48% | $0.40 | $2.55 | ||||||||||||||||||||||||||||||
| 6 | GLM 5.3 Flash | 47% | $0.07 | $0.25 | ||||||||||||||||||||||||||||||
| 7 | Kimi K3 | 46% | $2.85 | $14.25 | ||||||||||||||||||||||||||||||
| 8 | Qwen3.8-Flash-Next | 45% | n/a | n/a | ||||||||||||||||||||||||||||||
| 9 | GPT-5.6 Sol | 44% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 10 | Claude Opus 5 | 42% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 11 | Grok 4.5 | 42% | $2.00 | $6.00 | ||||||||||||||||||||||||||||||
| 12 | DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 41% | n/a | n/a | ||||||||||||||||||||||||||||||
| 13 | GPT-5.6 Terra | 40% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 14 | DeepSeek V4 Pro 0423 | 40% | $0.43 | $0.87 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 15 | GPT-5.4 | 40% | $2.50 | $15.00 | ||||||||||||||||||||||||||||||
| 16 | DeepSeek V4 Flash 0423 | 39% | $0.09 | $0.17 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 17 | GPT-5.5 | 39% | $5.00 | $30.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 18 | Claude Fable 5 | 38% | $10.00 | $50.00 | ||||||||||||||||||||||||||||||
| 19 | Claude Sonnet 5 | 37% | $2.00 | $10.00 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 20 | Agnes 2.5 Pro Beta | Sapiens-ai | 36% | n/a | n/a | |||||||||||||||||||||||||||||
| 21 | Motif 3 | Motif-technologies | 35% | n/a | n/a | |||||||||||||||||||||||||||||
| 22 | Muse Spark 1.2 | Other | 35% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 23 | GLM 5.2 | 35% | $0.61 | $1.98 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 24 | Claude Opus 4.7 | 35% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 25 | Claude Sonnet 4.6 | 34%reasoning: adaptive | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| 26 | Claude Opus 4.8 | 34% | $5.00 | $25.00 | ||||||||||||||||||||||||||||||
| 27 | Gemini 3.7 Flash | 33% | $0.38 | $1.88 | ||||||||||||||||||||||||||||||
| 28 | Gemini 3.5 Flash | 32% | $1.50 | $9.00 | ||||||||||||||||||||||||||||||
| 29 | Muse Spark 1.1 | Other | 32% | $1.25 | $4.25 | |||||||||||||||||||||||||||||
| 30 | GPT-5.6 Luna | 31% | $0.20 | $1.20 | ||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||
| 31 | DeepSeek V4 Flash (max) | 31% | n/a | n/a | ||||||||||||||||||||||||||||||
| 32 | DeepSeek V4 Pro (max) | 30% | n/a | n/a | ||||||||||||||||||||||||||||||
| 33 | Gemini 3.6 Flash | 30% | $0.75 | $3.75 | ||||||||||||||||||||||||||||||
| 34 | Inkling | Other | 29% | $0.95 | $4.05 | |||||||||||||||||||||||||||||
| 35 | Motif 3 (Beta) | Motif-technologies | 29% | n/a | n/a | |||||||||||||||||||||||||||||
| 36 | JT-4.1 Flash 236B A21B | China-mobile | 28% | n/a | n/a | |||||||||||||||||||||||||||||
| 37 | GPT-5.4 Nano | 27% | $0.20 | $1.25 | ||||||||||||||||||||||||||||||
| 38 | Ling-3.0-flash | Other | 27% | $0.02 | $0.06 | |||||||||||||||||||||||||||||
| 39 | GPT-5.4 Mini | 26% | $0.75 | $4.50 | ||||||||||||||||||||||||||||||
| 40 | Claude 4.5 Sonnet | 25%reasoning: true | $3.00 | $15.00 | ||||||||||||||||||||||||||||||
| 41 | Muse Glimmer | 24% | n/a | n/a | ||||||||||||||||||||||||||||||
| 42 | Kimi K2.6 | 23% | $0.65 | $3.40 | ||||||||||||||||||||||||||||||
| 43 | Solar Pro 4 | Other | 23% | $0.03 | $0.12 | |||||||||||||||||||||||||||||
| 44 | Hy3 | Other | 23% | $0.13 | $0.53 | |||||||||||||||||||||||||||||
| 45 | GPT-5 | 22% | $1.25 | $10.00 | ||||||||||||||||||||||||||||||
| 46 | G9v3-39A5B | Ai9stars | 22% | n/a | n/a | |||||||||||||||||||||||||||||
| 47 | Solar Open2 250B | Upstage | 22% | n/a | n/a | |||||||||||||||||||||||||||||
| 48 | Gemini 3.1 Pro Preview | 21% | $2.00 | $12.00 | ||||||||||||||||||||||||||||||
| 49 | DeepSeek V3.1 Terminus | 21%reasoning: true | $0.27 | $1.00 | ||||||||||||||||||||||||||||||
| 50 | Qwen3.6 Plus | 21% | $0.33 | $1.95 | ||||||||||||||||||||||||||||||
Frequently asked questions
What is the τ-bench Banking benchmark?
The banking domain in Sierra Research's tau-bench family: a knowledge-retrieval customer-service scenario where an agent answers banking questions using a retrieval pipeline rather than action APIs. Added in the τ³-bench update to the tau2-bench repository.
Which AI model scores highest on τ-bench Banking?
Qwen3.8 Max (Alibaba) leads with 51%, from $1.65 per 1M input tokens.
How many models are ranked on τ-bench Banking?
165 models carry a τ-bench Banking score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.
Other benchmarks
compare the same models on a different evalEvery weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI