Tau-Bench Retail — leaderboard
Agentic customer service benchmark in a simulated retail environment. Models handle multi-turn conversations with tool-calling to resolve customer requests. From Sierra AI (tau2-bench).
Metric: Pass@1 (%). Source: taubench.com. Status: saturation imminent. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 86.2 |
| 2 | Gemini 3 Pro | 85.3 |
| 3 | Claude Opus 4.1 | 82.4 |
| 4 | GPT-5 | 81.6 |
| 5 | Claude Opus 4 | 81.4 |
| 6 | DeepSeek V3.2 | 81.1 |
| 7 | Claude Sonnet 4 | 80.5 |
| 8 | Qwen 3 Max (Thinking) | 79.4 |
| 9 | GPT-4.1 | 74 |
| 10 | O3 | 73.9 |
| 11 | Qwen 3 Max | 72.2 |
| 12 | Claude 3.7 Sonnet | 72.1 |
| 13 | Kimi K2 | 70.6 |
| 14 | O4 Mini | 68.3 |
| 15 | GPT-4.1 Mini | 61.4 |
Interactive version: theaggregate.ai/benchmark?slug=tau-bench-retail · How the rankings work · Data refreshed daily, snapshot 2026-07-22.