AgentGym2 - Data Analysis: leaderboard

Metric: Task success (%; Avg@3 on the 57 end-to-end data analysis tasks, data cleaning included; mean over three runs of the share of tasks answered correctly (Avg@3), short string or numeric answers graded CORRECT by a Qwen3-235B-A22B-Instruct-2507 judge; ReAct agent with OpenAI tool schemas, tools to be discovered in a de-idealized environment with noisy or underspecified inputs, at most 100 interaction rounds, temperature 0.6, thinking mode as printed per model). Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)39.77
2GPT-539.18
3DeepSeek V3.1 (Non-reasoning)20.47
4Kimi K216.96
5DeepSeek V3.2 Exp (Non-reasoning)13.45
6Gemini 2.5 Pro12.28
7Qwen 3 235B A22B 2507 Instruct9.94
8GLM-4.68.19
9Qwen 3 235B A22B 2507 (Thinking)2.92
10Qwen 3 32B (Thinking)2.34
11Qwen 3 8B (Thinking)1.75
12Qwen 3 32B (Non-reasoning)1.75
13Qwen 3 8B (Non-reasoning)0

Interactive version: theaggregate.ai/benchmark?slug=agentgym2-data-analysis · How It Works · Data refreshed daily, snapshot 2026-09-29.