DSAgentBench (Screenshot + A11y Tree): leaderboard

Metric: Task success rate (%; percentage of 275 end-to-end data-science tasks (data acquisition, exploration, feature engineering, modeling, visualization, evaluation) whose final artifacts pass a deterministic Python evaluator (score >= 0.95), run in a real OSWorld desktop with VS Code, Jupyter and a terminal; temperature 0.1, 15-step budget; screenshot plus accessibility-tree observation). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.656.7
2GPT-529.81
3GPT-4o24.54
4Gemini 2.5 Pro20.81
5GPT-5 Mini19.03
6Claude Sonnet 4.59.21
7Claude Sonnet 44.64
8O4 Mini2.55

Interactive version: theaggregate.ai/benchmark?slug=dsagentbench-screenshot-plus-a11y-tree · How It Works · Data refreshed daily, snapshot 2026-09-26.