SCOPE Experimental Design (CoT-only): leaderboard
Metric: Total plan score (out of 30): sum of six sub-dimension rubric scores from a GPT-5.2 judge against the ground-truth experimental design of 300 recent ICML, NeurIPS and ICLR papers in 19 research domains; a redline rule zeros a sub-dimension with a fatal flaw such as a hallucinated dataset or baseline, an incompatible metric or a violated constraint; chain-of-thought prompting without search. Source: arxiv.org. Saturation forecast: Around February 2028. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 18.62 |
| 2 | GPT-5.2 | 18.22 |
| 3 | Kimi K2 | 15.62 |
| 4 | DeepSeek V3.2 | 13.92 |
| 5 | Qwen 3 Max | 13.05 |
| 6 | Gemini 3 Pro | 12.65 |
| 7 | Grok 4 | 12.55 |
Interactive version: theaggregate.ai/benchmark?slug=scope-experimental-design-cot-only · How It Works · Data refreshed daily, snapshot 2026-09-26.