SCOPE Experimental Design (CoT-only): leaderboard

Metric: Total plan score (out of 30): sum of six sub-dimension rubric scores from a GPT-5.2 judge against the ground-truth experimental design of 300 recent ICML, NeurIPS and ICLR papers in 19 research domains; a redline rule zeros a sub-dimension with a fatal flaw such as a hallucinated dataset or baseline, an incompatible metric or a violated constraint; chain-of-thought prompting without search. Source: arxiv.org. Saturation forecast: Around February 2028. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.518.62
2GPT-5.218.22
3Kimi K215.62
4DeepSeek V3.213.92
5Qwen 3 Max13.05
6Gemini 3 Pro12.65
7Grok 412.55

Interactive version: theaggregate.ai/benchmark?slug=scope-experimental-design-cot-only · How It Works · Data refreshed daily, snapshot 2026-09-26.