DARE-bench (Data Science) - Classification (Instruction Following): leaderboard

Metric: Strict accuracy (%) on the 74 classification instruction-following tasks: the agent must reproduce a reference workflow exactly and scores 1 only when its predictions equal the reference solution's; DARE-bench test tasks derived from recently updated Kaggle datasets; the model works as a data-science agent with a sandboxed Python execution tool (5 interaction turns, 200 s per execution, greedy decoding), mean of three repeats; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-569.81#91
2O4 Mini67.56#172
3Claude 3.7 Sonnet61.48#241
4GPT-4.155.82#240
5GPT-4o32.88#333
6Qwen 3 32B17.11#424
7Claude Sonnet 416.21#194
8Qwen 3 4B3.6#823

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=dare-bench-data-science-classification-instruction-following · How It Works · Data refreshed daily, snapshot 2026-10-11.