BenchBench-Protocol (Best of 10): leaderboard

Metric: Expected best-of-10 normalized rubric score (%; per task the best of ten sampled responses by weighted rubric score, 149 protocol-modification tasks, graded by GPT-5.6 Terra). Source: arxiv.org. Saturation forecast: Around July 2027. 9 models tracked.

Top models

#ModelScore
1Claude Opus 5 (Max)75.3
2Kimi K3 (Max)62.8
3GPT-5.6 Terra (Max)57.7
4GPT-5.6 Sol (Max)57.5
5GLM-5.2 (Max)54.9
6Inkling (Max)54.1
7Grok 4.5 (High)53.3
8GPT-5.6 Luna (Max)52.7
9Gemini 3.6 Flash (High)47.6

Interactive version: theaggregate.ai/benchmark?slug=benchbench-protocol-best-of-10 · How It Works · Data refreshed daily, snapshot 2026-09-26.