τ²-bench benchmark

314 AI models ranked on τ²-bench. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

GLM 5.2 (Zhipu) leads at 99%, from $0.61 per 1M input.

Data via GitHub: sierra-research/tau2-bench

Models ranked

314

on this benchmark

Top score

99%

current leader

Type

Single eval

benchmark

Updated

30 Aug 2026

last source fetch

Cost vs τ²-bench

193 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on τ²-bench

highest score first · 314 models
Maker
#
1GLM 5.2Zhipu99%$0.61$1.98
2JT-35B-FlashChina-mobile99%n/an/a
3GLM 4.7 FlashZhipu99%$0.00$0.00
4Claude Fable 5Anthropic99%$10.00$50.00
5Step 3.7 FlashOther99%$0.20$1.15
6GLM 5 TurboZhipu99%$1.20$4.00
7GLM 5V TurboZhipu99%$1.20$4.00
8GLM 5Zhipu98%$0.60$1.92
9GLM 5.1Zhipu98%$0.97$3.04
10Grok 4.3xAI98%$1.25$2.50
11Qwen3.6 PlusAlibaba98%$0.33$1.95
12Grok 4.20 0309Spacexai96%n/an/a
13DeepSeek V4 Pro 0423DeepSeek96%$0.43$0.87
14DeepSeek V4 Pro (max)DeepSeek96%n/an/a
15Kimi K2.5Moonshot96%$0.45$2.25
16Kimi K2.6Moonshot96%$0.65$3.40
17GLM 4.7Zhipu96%$0.40$1.75
18Qwen3.6 Max PreviewAlibaba96%n/an/a
19DeepSeek V4 Flash 0423DeepSeek96%high$0.08$0.17
20Qwen3.5 397B A17BAlibaba96%$0.39$2.34
21Gemini 3.5 FlashGoogle96%medium$1.50$9.00
22Gemini 3.1 Pro PreviewGoogle96%$2.00$12.00
23Qwen3.6 35B A3BAlibaba95%$0.10$0.45
24MiniMax M2.5MiniMax95%$0.27$1.08
25Mimo V2 FlashXiaomi95%reasoning: true$0.10$0.30
26DeepSeek V4 Flash (max)DeepSeek95%n/an/a
27MiMo-V2-ProXiaomi95%n/an/a
28Qwen3.7 MaxAlibaba95%$1.25$3.75
29Claude Opus 4.8Anthropic94%$5.00$25.00
30Step 3.5 FlashStepfun94%n/an/a
31Qwen3.6 27BAlibaba94%$0.15$0.50
32MiMo-V2.5-ProXiaomi94%$0.43$0.87
33Mistral Medium 3.5Mistral94%$1.50$7.50
34Qwen3.5-27BAlibaba94%$0.20$1.56
35GPT-5.5OpenAI94%$5.00$30.00
36Qwen3.5-122B-A10BAlibaba94%$0.25$1.75
37Grok 4.1 Fast ReasoningxAI93%$0.20$0.50
38MiMo-V2-Flash (Feb 2026)Xiaomi93%n/an/a
39Tri-21B-think PreviewTrillion-labs93%n/an/a
40Kimi K2 ThinkingMoonshot93%$0.60$1.20
41Qwen3.7 PlusAlibaba93%$0.32$1.28
42Grok 4.20xAI93%$1.25$2.50
43JT-MINIChina-mobile93%n/an/a
44Hy3 previewOther93%$0.18$0.60
45Nova 2.0 Pro PreviewAmazon93%mediumn/an/a
46Hy3Other93%$0.13$0.53
47Ring-2.6-1TOther92%$0.07$0.63
48Claude Opus 4.6Anthropic92%reasoning: adaptive$5.00$25.00
49GPT-5.2-CodexOpenAI92%$1.75$14.00
50Qwen3.5 4BAlibaba92%n/an/a
150 of 314
Page 1 of 7
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the τ²-bench benchmark?

Sierra Research's τ²-bench evaluates LLM agents in realistic multi-turn customer-service scenarios (airline, retail, telecom) where both the agent and a simulated user take actions using tools under domain policies. Scored on task success (final vs goal state) and on pass^k reliability across repeated attempts.

Which AI model scores highest on τ²-bench?

GLM 5.2 (Zhipu) leads with 99%, from $0.61 per 1M input tokens.

How many models are ranked on τ²-bench?

314 models carry a τ²-bench score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 30 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI