SlopCodeBench (Just Solve) - Core: leaderboard
Metric: Core solve rate (%): share of the 196 checkpoints whose workspace passes the core tests (functionality stated or shown in the specification), on SlopCodeBench's 36 iterative problems (196 checkpoints): the agent extends its own prior code as the CLI or API specification evolves, with hidden tests, run in the model's native coding harness (the harness and reasoning level are in the label), using the plain task prompt; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Claude Opus 4.6 (High) | 67.3 | #60 (Claude Opus 4.6) |
| 2 | GPT-5.5 (High) | 66.8 | #26 (GPT-5.5) |
| 3 | Claude Opus 4.7 (High) | 65.8 | #45 (Claude Opus 4.7) |
| 4 | GPT-5.4 (High) | 62.8 | #76 (GPT-5.4) |
| 5 | GPT-5.3 Codex (High) | 60.7 | #68 (GPT-5.3 Codex) |
| 6 | Claude Sonnet 4.6 (High) | 57.7 | #85 (Claude Sonnet 4.6) |
| 7 | Claude Opus 4.5 (High) | 57.7 | #79 (Claude Opus 4.5) |
| 8 | GPT-5.2 Codex (High) | 56.1 | #89 (GPT-5.2 Codex) |
| 9 | GPT-5.4 Mini (High) | 53.1 | #202 (GPT-5.4 Mini) |
| 10 | Kimi K2.5 | 31.6 | #139 |
| 11 | Kimi K2.5 (High) | 29.6 | #139 (Kimi K2.5) |
| 12 | MiniMax-M2.7 | 21.9 | #248 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=slopcodebench-just-solve-core · How It Works · Data refreshed daily, snapshot 2026-10-11.