GeoNatureAgent Benchmark: leaderboard

Metric: Task accuracy (%), mean of three seeds, on 93 environmental geospatial analysis tasks in 18 capability categories (tool selection, spatial reasoning, multi-turn memory, error handling, multilingual queries); a ReAct agent calls 16 tools on an MCP geospatial API serving three indicators for Spain and Portugal, and a task passes on the expected tool calls and required answer strings; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Claude Sonnet 460.8
2DeepSeek V3.256.3
3GLM-550.2
4Gemini 2.5 Pro48
5GPT-4o41.6
6Qwen 3 235B A22B41.2
7GPT-OSS-120B34.1
8Llama 4 Scout26.9
9Gemma 3 27B15.8

Interactive version: theaggregate.ai/benchmark?slug=geonatureagent-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-29.