SCOPE Experimental Design - Baselines (CoT-only): leaderboard

Metric: Sub-dimension score (out of 5) for baseline selection: rubric scores from a GPT-5.2 judge against the ground-truth experimental design of 300 recent ICML, NeurIPS and ICLR papers in 19 research domains; a redline rule zeros a sub-dimension with a fatal flaw such as a hallucinated dataset or baseline, an incompatible metric or a violated constraint; chain-of-thought prompting without search. Source: arxiv.org. Saturation forecast: Around 2029. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.52.27
2GPT-5.22.15
3DeepSeek V3.21.94
4Kimi K21.89
5Qwen 3 Max1.75
6Gemini 3 Pro1.72
7Grok 41.59

Interactive version: theaggregate.ai/benchmark?slug=scope-experimental-design-baselines-cot-only · How It Works · Data refreshed daily, snapshot 2026-09-26.