AgentNoiseBench (tau2-bench, Tool Noise): leaderboard

Metric: Success on a stratified 25 percent sample of tau2-bench tasks (airline and retail) with tool noise injected into tool returns (execution failures, incomplete, erroneous, misleading or redundant outputs); noise produced by a fixed generator whose prompt was optimized to degrade a reference agent while keeping every task solvable; task success rate (fraction times 100), mean of four trials; the paper's stability-gated criterion counts a task solved only with a correct final outcome and no noise-induced deviation in the trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 24 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 (Thinking)61#105 (GPT-5.2)
2Claude Sonnet 4.5 (Thinking)57#138 (Claude Sonnet 4.5)
3O357#121
4Qwen 3 Max (Thinking)56#201 (Qwen 3 Max)
5GLM-4.6 (Thinking)56#246 (GLM-4.6)
6GLM-4.5 (Thinking)49#265 (GLM-4.5)
7Claude Sonnet 4.548#138
8GLM-4.6 (Non-reasoning)48#246 (GLM-4.6)
9LongCat-Flash-Chat48#322
10Claude Sonnet 4 (Thinking)46#194 (Claude Sonnet 4)
11GLM-4.5 (Non-reasoning)42#265 (GLM-4.5)
12Claude Sonnet 441#194
13Doubao-Seed-1.6 (Thinking)40
14Gemini 2.5 Pro37#145
15DeepSeek R1 052837#217

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agentnoisebench-tau2-bench-tool-noise · How It Works · Data refreshed daily, snapshot 2026-10-11.