SEATauBench (L2 Tool) - Thai Retail: leaderboard

Metric: pass@1 (%; mean success over three trials per task of reaching the expected final state; the 114 tau2-bench retail tasks with the tool schemas shown to the agent are in Thai; dialogue, policy and database stay in English; Qwen3-235B-A22B-Instruct-2507 user simulator, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.

Top models

#ModelScore
1GPT-5 Mini64.6
2Kimi K2.560.2
3Qwen 3 235B A22B 2507 Instruct56.1

Interactive version: theaggregate.ai/benchmark?slug=seataubench-l2-tool-thai-retail · How It Works · Data refreshed daily, snapshot 2026-09-26.