DABstep — leaderboard

Data Agent Benchmark for multi-step reasoning by Adyen: ~450 data analysis questions across easy and hard levels requiring agents to analyze CSVs and cross-reference documents.

Metric: Hard Level Accuracy (%). Source: huggingface.co. Status: saturation imminent. 28 models tracked.

Top models

#ModelScore
1Claude Haiku 4.589.95
2GPT-557.67
3Gemini 2.5 Pro45.24
4Gemini 2.5 Pro (Preview 05-06)41.01
5Claude 3.5 Sonnet (20241022)28.04
6Claude Sonnet 4 (20250514)19.84
7DeepSeek V316.4
8O4 Mini14.55
9Claude 3.7 Sonnet (20250219)13.76
10O3 Mini13.76
11GPT-4.112.43
12O111.11
13Gemini 2.0 Flash9.79
14Claude 3.5 Sonnet9.26
15GPT-4.1 Mini8.99

Interactive version: theaggregate.ai/benchmark?slug=dabstep · How the rankings work · Data refreshed daily, snapshot 2026-07-22.