τ-bench Banking benchmark

165 AI models ranked on τ-bench Banking. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Qwen3.8 Max (Alibaba) leads at 51%, from $1.65 per 1M input.

Data via GitHub: sierra-research/tau2-bench

Models ranked

165

on this benchmark

Top score

51%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs τ-bench Banking

111 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on τ-bench Banking

highest score first · 165 models
Maker
#
1Qwen3.8 MaxAlibaba51%$1.65$4.95
2Grok 4.6xAI51%$2.00$6.00
3GLM 5.3Zhipu50%$1.40$4.40
4Qwen3.8 2.4T A95BAlibaba49%$2.00$6.00
5Qwen3.8 27BAlibaba48%$0.40$2.55
6GLM 5.3 FlashZhipu47%$0.07$0.25
7Kimi K3Moonshot46%$2.85$14.25
8Qwen3.8-Flash-NextAlibaba45%n/an/a
9GPT-5.6 SolOpenAI44%$2.00$10.00
10Claude Opus 5Anthropic42%$5.00$25.00
11Grok 4.5xAI42%$2.00$6.00
12DeepSeek V4 Flash Vision (Reasoning, Max Effort)DeepSeek41%n/an/a
13GPT-5.6 TerraOpenAI40%$2.00$12.00
14DeepSeek V4 Pro 0423DeepSeek40%$0.43$0.87
15GPT-5.4OpenAI40%$2.50$15.00
16DeepSeek V4 Flash 0423DeepSeek39%$0.09$0.17
17GPT-5.5OpenAI39%$5.00$30.00
18Claude Fable 5Anthropic38%$10.00$50.00
19Claude Sonnet 5Anthropic37%$2.00$10.00
20Agnes 2.5 Pro BetaSapiens-ai36%n/an/a
21Motif 3Motif-technologies35%n/an/a
22Muse Spark 1.2Other35%$1.25$4.25
23GLM 5.2Zhipu35%$0.61$1.98
24Claude Opus 4.7Anthropic35%$5.00$25.00
25Claude Sonnet 4.6Anthropic34%reasoning: adaptive$3.00$15.00
26Claude Opus 4.8Anthropic34%$5.00$25.00
27Gemini 3.7 FlashGoogle33%$0.38$1.88
28Gemini 3.5 FlashGoogle32%$1.50$9.00
29Muse Spark 1.1Other32%$1.25$4.25
30GPT-5.6 LunaOpenAI31%$0.20$1.20
31DeepSeek V4 Flash (max)DeepSeek31%n/an/a
32DeepSeek V4 Pro (max)DeepSeek30%n/an/a
33Gemini 3.6 FlashGoogle30%$0.75$3.75
34InklingOther29%$0.95$4.05
35Motif 3 (Beta)Motif-technologies29%n/an/a
36JT-4.1 Flash 236B A21BChina-mobile28%n/an/a
37GPT-5.4 NanoOpenAI27%$0.20$1.25
38Ling-3.0-flashOther27%$0.02$0.06
39GPT-5.4 MiniOpenAI26%$0.75$4.50
40Claude 4.5 SonnetAnthropic25%reasoning: true$3.00$15.00
41Muse GlimmerMeta24%n/an/a
42Kimi K2.6Moonshot23%$0.65$3.40
43Solar Pro 4Other23%$0.03$0.12
44Hy3Other23%$0.13$0.53
45GPT-5OpenAI22%$1.25$10.00
46G9v3-39A5BAi9stars22%n/an/a
47Solar Open2 250BUpstage22%n/an/a
48Gemini 3.1 Pro PreviewGoogle21%$2.00$12.00
49DeepSeek V3.1 TerminusDeepSeek21%reasoning: true$0.27$1.00
50Qwen3.6 PlusAlibaba21%$0.33$1.95
150 of 165
Page 1 of 4
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the τ-bench Banking benchmark?

The banking domain in Sierra Research's tau-bench family: a knowledge-retrieval customer-service scenario where an agent answers banking questions using a retrieval pipeline rather than action APIs. Added in the τ³-bench update to the tau2-bench repository.

Which AI model scores highest on τ-bench Banking?

Qwen3.8 Max (Alibaba) leads with 51%, from $1.65 per 1M input tokens.

How many models are ranked on τ-bench Banking?

165 models carry a τ-bench Banking score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI