SEATauBench (L2 Interaction) - Filipino Telecom: leaderboard
Metric: pass@1 (%; mean success over three trials per task of reaching the expected final state; the 114 tau2-bench telecom tasks with the simulated user and the agent converse in Filipino; tools, policy and database stay in English; Qwen3-235B-A22B-Instruct-2507 user simulator, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 88.3 |
| 2 | GPT-5 Mini | 74 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 15.5 |
Interactive version: theaggregate.ai/benchmark?slug=seataubench-l2-interaction-filipino-telecom · How It Works · Data refreshed daily, snapshot 2026-09-26.