DSAgentBench (Screenshot): leaderboard
Metric: Task success rate (%; percentage of 275 end-to-end data-science tasks (data acquisition, exploration, feature engineering, modeling, visualization, evaluation) whose final artifacts pass a deterministic Python evaluator (score >= 0.95), run in a real OSWorld desktop with VS Code, Jupyter and a terminal; temperature 0.1, 15-step budget; screenshot-only observation). Source: arxiv.org. Saturation forecast: Around February 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 50.55 |
| 2 | GPT-5 | 23.63 |
| 3 | GPT-4o | 19.34 |
| 4 | GPT-5 Mini | 15.2 |
| 5 | Gemini 2.5 Pro | 14.49 |
| 6 | Claude Sonnet 4 | 4.55 |
| 7 | Claude Sonnet 4.5 | 3.17 |
| 8 | O4 Mini | 1.82 |
Interactive version: theaggregate.ai/benchmark?slug=dsagentbench-screenshot · How It Works · Data refreshed daily, snapshot 2026-09-26.