Toolathlon: leaderboard

Tool-use benchmark spanning many tool categories, testing whether agents can select, sequence, and combine tools to complete realistic tasks.

Metric: Score (self-reported). Source: benchmarklist.com. Status: years away from saturation. 47 models tracked.

Top models

#ModelScore
1Claude Opus 5 (Max)80.6
2Claude Opus 4.8 (Max)79.9
3GLM-5.3 Flash78.4
4Claude Fable 5.1 (Max)77.8
5Claude Opus 577.6
6Kimi K376.5
7Muse Spark 1.175.6
8GPT-5.6 Sol74.9
9GPT-5.6 Terra74.9
10Claude Sonnet 5 (Max)74.7
11GPT-5.6 Luna74.6
12DeepSeek V4 Pro (0813)74.1
13GPT-5.5 (xHigh)73.5
14Qwen3.8-Flash-Next73.5
15GLM-5.373

Interactive version: theaggregate.ai/benchmark?slug=toolathlon · How It Works · Data refreshed daily, snapshot 2026-09-05.