GTA — leaderboard

General Tool Agents benchmark from OpenCompass. Tests LLMs on multi-step tool use across perception, operation, logic, and creativity categories using real-world tool chains.

Metric: Answer Accuracy (%). Source: github.com. Status: saturation imminent. 13 models tracked.

Top models

#ModelScore
1GPT-545.07
2DeepSeek V3.242.68
3Qwen 3 235B A22B42.21
4Gemini 2.5 Pro39.58
5Qwen 3 8B27.1
6Llama 3.2 3B Instruct24.99
7Llama 4 Scout22.03
8Kimi K218.32
9Claude Sonnet 4.517.89
10Grok 417.78
11Qwen 3 30B A3B15.67
12Llama 3.1 70B Instruct13.24
13Llama 3.1 8B Instruct8.78

Interactive version: theaggregate.ai/benchmark?slug=gta · How the rankings work · Data refreshed daily, snapshot 2026-07-22.