TerraBench - NumScore: leaderboard

Metric: NumScore (%): tolerance-aware partial credit for numeric answers, halving for each extra tolerance width of error (missing or unparseable answers score 0), on the 403 executable TerraBench tasks (Fundamentals, Simulator-Grounded and Document-Grounded Verification tracks across eight Earth-science domains), each model driving the benchmark's TerraAgent ReAct harness with its 77 scientific tools; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 17 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.628.44
2GPT-5.526.45
3Claude Haiku 4.521.19
4Gemini 3.1 Pro (Preview)17.07
5GPT-5.412.29
6Gemini 2.5 Flash7.7
7Qwen 3.5 35B A3B7.49
8Gemma 4 26B A4B4.45
9Qwen 3 14B4.3
10Gemma 4 E4B3.04
11Mistral 7B Instruct (v0.3)2.41
12Qwen 3 8B2.06
13InternVL3-8B1.7
14Qwen 3.5 9B1.3
15Qwen 3 1.7B1.15

Interactive version: theaggregate.ai/benchmark?slug=terrabench-numscore · How It Works · Data refreshed daily, snapshot 2026-09-29.