RedundancyBench (Window-to-One): leaderboard
Metric: Trajectory-level accuracy (%; mean over the airline, retail and telecom domains of the share of successful tau2-bench trajectories (Qwen-3.6-Plus agent, synthetic redundant steps inserted, three-round expert annotation) correctly classified as containing a redundant step or not; Window-to-One: the LLM judges each step with the three preceding and three following steps as context). Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 (Non-reasoning) | 70.88 |
| 2 | DeepSeek V4 Pro (Thinking) | 68.48 |
| 3 | GPT-4o | 64.81 |
Interactive version: theaggregate.ai/benchmark?slug=redundancybench-window-to-one · How It Works · Data refreshed daily, snapshot 2026-09-26.