TerraBench - ToolUseScore: leaderboard

Metric: ToolUseScore (%): weighted composite of instruction validity, tool-call success, tool-selection accuracy, tool-category F1, argument accuracy and workflow-order consistency against the canonical trace, on the 403 executable TerraBench tasks (Fundamentals, Simulator-Grounded and Document-Grounded Verification tracks across eight Earth-science domains), each model driving the benchmark's TerraAgent ReAct harness with its 77 scientific tools; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 17 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.659.22
2Claude Haiku 4.548.14
3GPT-5.545.01
4Gemini 3.1 Pro (Preview)44.93
5Gemma 4 26B A4B41.84
6GPT-5.441.09
7Qwen 3.5 35B A3B39.95
8Gemini 2.5 Flash36.2
9Qwen 3.5 9B31.18
10Qwen 3 14B21.12
11Mistral 7B Instruct (v0.3)18.79
12Gemma 4 E4B16.31
13Qwen 3 8B7.97
14Qwen 3 1.7B5.4
15InternVL3-8B4.6

Interactive version: theaggregate.ai/benchmark?slug=terrabench-toolusescore · How It Works · Data refreshed daily, snapshot 2026-09-29.