Toolathlon — leaderboard

Tool-use benchmark spanning many tool categories, testing whether agents can select, sequence, and combine tools to complete realistic tasks.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 18 models tracked.

Top models

#ModelScore
1Claude Fable 561.7
2Claude Mythos 561.7
3Claude Mythos Preview61.1
4Claude Opus 4.859.9
5Claude Opus 4.759.3
6Claude Opus 4.656.8
7GPT-5.555.6
8GPT-5.454.6
9Claude Sonnet 554.3
10DeepSeek V4 Pro52.8
11Nex N2 Pro51.9
12Claude Sonnet 4.649.4
13Gemini 3.1 Pro (Preview)48.8
14GLM-5.248.2
15Claude Sonnet 4.541

Interactive version: theaggregate.ai/benchmark?slug=toolathlon · How the rankings work · Data refreshed daily, snapshot 2026-07-22.