DSAgentBench (Screenshot + A11y Tree): leaderboard
Metric: Task success rate (%; percentage of 275 end-to-end data-science tasks (data acquisition, exploration, feature engineering, modeling, visualization, evaluation) whose final artifacts pass a deterministic Python evaluator (score >= 0.95), run in a real OSWorld desktop with VS Code, Jupyter and a terminal; temperature 0.1, 15-step budget; screenshot plus accessibility-tree observation). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 56.7 |
| 2 | GPT-5 | 29.81 |
| 3 | GPT-4o | 24.54 |
| 4 | Gemini 2.5 Pro | 20.81 |
| 5 | GPT-5 Mini | 19.03 |
| 6 | Claude Sonnet 4.5 | 9.21 |
| 7 | Claude Sonnet 4 | 4.64 |
| 8 | O4 Mini | 2.55 |
Interactive version: theaggregate.ai/benchmark?slug=dsagentbench-screenshot-plus-a11y-tree · How It Works · Data refreshed daily, snapshot 2026-09-26.