AgentNoiseBench (VitaBench, User Noise): leaderboard

Metric: Success on a stratified 25 percent sample of VitaBench tasks (delivery and in-store) with user noise injected into the simulated user's turns (ambiguous, inconsistent, redundant, topic-drifting or boundary-probing instructions); noise produced by a fixed generator whose prompt was optimized to degrade a reference agent while keeping every task solvable; task success rate (fraction times 100), mean of four trials; the paper's stability-gated criterion counts a task solved only with a correct final outcome and no noise-induced deviation in the trajectory; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 24 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 (Thinking)45#105 (GPT-5.2)
2O344#121
3Claude Sonnet 4.5 (Thinking)41#138 (Claude Sonnet 4.5)
4Qwen 3 Max (Thinking)40#201 (Qwen 3 Max)
5GLM-4.6 (Thinking)40#246 (GLM-4.6)
6Claude Sonnet 4.539#138
7Gemini 2.5 Pro38#145
8GLM-4.5 (Thinking)37#265 (GLM-4.5)
9DeepSeek R1 052836#217
10Claude Sonnet 4 (Thinking)36#194 (Claude Sonnet 4)
11LongCat-Flash-Chat36#322
12GPT-4.133#240
13Claude Sonnet 433#194
14GLM-4.6 (Non-reasoning)33#246 (GLM-4.6)
15DeepSeek V3.2 Exp (Non-reasoning)32#227 (DeepSeek V3.2 Exp)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agentnoisebench-vitabench-user-noise · How It Works · Data refreshed daily, snapshot 2026-10-11.