SCOPE Experimental Design (CoT + Search): leaderboard

Metric: Total plan score (out of 30): sum of six sub-dimension rubric scores from a GPT-5.2 judge against the ground-truth experimental design of 300 recent ICML, NeurIPS and ICLR papers in 19 research domains; a redline rule zeros a sub-dimension with a fatal flaw such as a hallucinated dataset or baseline, an incompatible metric or a violated constraint; chain-of-thought prompting with web search the model is told to call when needed. Source: arxiv.org. Saturation forecast: Around July 2028. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.518.38
2GPT-5.216.77
3Kimi K215.6
4DeepSeek V3.213.76
5Grok 413.12
6Qwen 3 Max12.88
7Gemini 3 Pro12.66

Interactive version: theaggregate.ai/benchmark?slug=scope-experimental-design-cot-plus-search · How It Works · Data refreshed daily, snapshot 2026-09-26.