SlopCodeBench (Anti-Slop) - Isolated: leaderboard
Metric: Isolated solve rate (%): share of the 196 checkpoints whose workspace passes all of that checkpoint's own tests, ignoring regression tests, on SlopCodeBench's 36 iterative problems (196 checkpoints): the agent extends its own prior code as the CLI or API specification evolves, with hidden tests, run in the model's native coding harness (the harness and reasoning level are in the label), using a prompt listing verbosity and over-engineering patterns to avoid; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.4 (High) | 25.5 | #76 (GPT-5.4) |
| 2 | GPT-5.5 (High) | 24 | #26 (GPT-5.5) |
| 3 | GPT-5.3 Codex (High) | 23.5 | #68 (GPT-5.3 Codex) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=slopcodebench-anti-slop-isolated · How It Works · Data refreshed daily, snapshot 2026-10-11.