TerraBench: leaderboard

Metric: Hit@tol (%): share of answer fields within the field's tolerance band (exact match for strings and booleans), averaged over multi-field outputs, on the 403 executable TerraBench tasks (Fundamentals, Simulator-Grounded and Document-Grounded Verification tracks across eight Earth-science domains), each model driving the benchmark's TerraAgent ReAct harness with its 77 scientific tools; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 17 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.622.88
2GPT-5.521.03
3Claude Haiku 4.515.23
4Gemini 3.1 Pro (Preview)13.21
5GPT-5.49.8
6Gemini 2.5 Flash6.34
7Qwen 3.5 35B A3B5.89
8Qwen 3 14B3.8
9Gemma 4 26B A4B3.67
10Gemma 4 E4B2.9
11Mistral 7B Instruct (v0.3)1.67
12Qwen 3 8B1.33
13Qwen 3.5 9B1.2
14InternVL3-8B1.1
15Qwen 3 1.7B1.08

Interactive version: theaggregate.ai/benchmark?slug=terrabench · How It Works · Data refreshed daily, snapshot 2026-09-29.