DAREBench - Text - Multi-Step with Tools: leaderboard

Metric: Accuracy (%). Source: arxiv.org. 35 models tracked.

Top models

#ModelScore
1Claude Opus 4.865.3
2MiMo-V2.5-Pro61.2
3Claude Sonnet 4.659.9
4GPT-5.559.8
5GPT-5.459.6
6MiMo-V2-Pro58.4
7Claude Opus 4.658.3
8MiMo-V2.557.5
9Gemini 3.1 Pro (Preview)56.1
10Qwen 3.6 Plus54.8
11DeepSeek V4 Pro54.3
12Qwen 3.5 Plus53.9
13GLM-553.3
14MiniMax-M2.753
15Gemini 3.5 Flash51.7

Interactive version: theaggregate.ai/benchmark?slug=darebench-text-multi-step-with-tools · How It Works · Data refreshed daily, snapshot 2026-09-19.