BenchBench-Protocol (Best of 10): leaderboard
Metric: Expected best-of-10 normalized rubric score (%; per task the best of ten sampled responses by weighted rubric score, 149 protocol-modification tasks, graded by GPT-5.6 Terra). Source: arxiv.org. Saturation forecast: Around July 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 5 (Max) | 75.3 |
| 2 | Kimi K3 (Max) | 62.8 |
| 3 | GPT-5.6 Terra (Max) | 57.7 |
| 4 | GPT-5.6 Sol (Max) | 57.5 |
| 5 | GLM-5.2 (Max) | 54.9 |
| 6 | Inkling (Max) | 54.1 |
| 7 | Grok 4.5 (High) | 53.3 |
| 8 | GPT-5.6 Luna (Max) | 52.7 |
| 9 | Gemini 3.6 Flash (High) | 47.6 |
Interactive version: theaggregate.ai/benchmark?slug=benchbench-protocol-best-of-10 · How It Works · Data refreshed daily, snapshot 2026-09-26.