DSAgentBench (Screenshot): leaderboard

Metric: Task success rate (%; percentage of 275 end-to-end data-science tasks (data acquisition, exploration, feature engineering, modeling, visualization, evaluation) whose final artifacts pass a deterministic Python evaluator (score >= 0.95), run in a real OSWorld desktop with VS Code, Jupyter and a terminal; temperature 0.1, 15-step budget; screenshot-only observation). Source: arxiv.org. Saturation forecast: Around February 2027. 9 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.650.55
2GPT-523.63
3GPT-4o19.34
4GPT-5 Mini15.2
5Gemini 2.5 Pro14.49
6Claude Sonnet 44.55
7Claude Sonnet 4.53.17
8O4 Mini1.82

Interactive version: theaggregate.ai/benchmark?slug=dsagentbench-screenshot · How It Works · Data refreshed daily, snapshot 2026-09-26.