DARE-bench (Data Science) - Regression (Instruction Following): leaderboard

Metric: Strict accuracy (%) on the 45 regression instruction-following tasks: the agent must reproduce a reference workflow exactly and scores 1 only when its predictions equal the reference solution's; DARE-bench test tasks derived from recently updated Kaggle datasets; the model works as a data-science agent with a sandboxed Python execution tool (5 interaction turns, 200 s per execution, greedy decoding), mean of three repeats; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-557.24#91
2O4 Mini53.62#172
3GPT-4.152.17#240
4Claude 3.7 Sonnet46.37#241
5GPT-4o20.28#333
6Qwen 3 32B15.21#424
7Claude Sonnet 415.21#194
8Qwen 3 4B0.72#823

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=dare-bench-data-science-regression-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-11.