Instruction Stacking Collapse - Code: leaderboard
Metric: Mean follow rate (%; share of the 20 stacked instructions a response satisfies on HumanEval code prompts, judged by per-instruction verifiers; stacks drawn uniformly at random from 22 deterministic instructions in six categories, 30 stacks x 10 items per size, temperature 0). Source: arxiv.org. Saturation forecast: Around June 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 44.7 |
| 2 | GPT-5 Mini | 44.3 |
| 3 | Gemini 2.5 Flash | 34.4 |
Interactive version: theaggregate.ai/benchmark?slug=instruction-stacking-collapse-code · How It Works · Data refreshed daily, snapshot 2026-09-29.