RedundancyBench (Window-to-One): leaderboard

Metric: Trajectory-level accuracy (%; mean over the airline, retail and telecom domains of the share of successful tau2-bench trajectories (Qwen-3.6-Plus agent, synthetic redundant steps inserted, three-round expert annotation) correctly classified as containing a redundant step or not; Window-to-One: the LLM judges each step with the three preceding and three following steps as context). Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)70.88
2DeepSeek V4 Pro (Thinking)68.48
3GPT-4o64.81

Interactive version: theaggregate.ai/benchmark?slug=redundancybench-window-to-one · How It Works · Data refreshed daily, snapshot 2026-09-26.