AgentNoiseBench (tau2-bench, User Noise): leaderboard

Metric: Success on a stratified 25 percent sample of tau2-bench tasks (airline and retail) with user noise injected into the simulated user's turns (ambiguous, inconsistent, redundant, topic-drifting or boundary-probing instructions); noise produced by a fixed generator whose prompt was optimized to degrade a reference agent while keeping every task solvable; task success rate (fraction times 100), mean of four trials; the paper's stability-gated criterion counts a task solved only with a correct final outcome and no noise-induced deviation in the trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 24 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 Max (Thinking)78#201 (Qwen 3 Max)
2GPT-5.2 (Thinking)72#105 (GPT-5.2)
3O368#121
4GLM-4.6 (Thinking)68#246 (GLM-4.6)
5Claude Sonnet 4.5 (Thinking)65#138 (Claude Sonnet 4.5)
6Claude Sonnet 4.564#138
7GLM-4.6 (Non-reasoning)64#246 (GLM-4.6)
8GLM-4.5 (Thinking)63#265 (GLM-4.5)
9LongCat-Flash-Chat62#322
10Doubao-Seed-1.6 (Thinking)61
11Claude Sonnet 4 (Thinking)58#194 (Claude Sonnet 4)
12Gemini 2.5 Pro55#145
13Claude Sonnet 455#194
14GLM-4.5 (Non-reasoning)55#265 (GLM-4.5)
15Qwen 3 Max51#201

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agentnoisebench-tau2-bench-user-noise · How It Works · Data refreshed daily, snapshot 2026-10-11.