AgentGym2 - Data Analysis: leaderboard
Metric: Task success (%; Avg@3 on the 57 end-to-end data analysis tasks, data cleaning included; mean over three runs of the share of tasks answered correctly (Avg@3), short string or numeric answers graded CORRECT by a Qwen3-235B-A22B-Instruct-2507 judge; ReAct agent with OpenAI tool schemas, tools to be discovered in a de-idealized environment with noisy or underspecified inputs, at most 100 interaction rounds, temperature 0.6, thinking mode as printed per model). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 (Thinking) | 39.77 |
| 2 | GPT-5 | 39.18 |
| 3 | DeepSeek V3.1 (Non-reasoning) | 20.47 |
| 4 | Kimi K2 | 16.96 |
| 5 | DeepSeek V3.2 Exp (Non-reasoning) | 13.45 |
| 6 | Gemini 2.5 Pro | 12.28 |
| 7 | Qwen 3 235B A22B 2507 Instruct | 9.94 |
| 8 | GLM-4.6 | 8.19 |
| 9 | Qwen 3 235B A22B 2507 (Thinking) | 2.92 |
| 10 | Qwen 3 32B (Thinking) | 2.34 |
| 11 | Qwen 3 8B (Thinking) | 1.75 |
| 12 | Qwen 3 32B (Non-reasoning) | 1.75 |
| 13 | Qwen 3 8B (Non-reasoning) | 0 |
Interactive version: theaggregate.ai/benchmark?slug=agentgym2-data-analysis · How It Works · Data refreshed daily, snapshot 2026-09-29.