GeoNatureAgent Benchmark: leaderboard
Metric: Task accuracy (%), mean of three seeds, on 93 environmental geospatial analysis tasks in 18 capability categories (tool selection, spatial reasoning, multi-turn memory, error handling, multilingual queries); a ReAct agent calls 16 tools on an MCP geospatial API serving three indicators for Spain and Portugal, and a task passes on the expected tool calls and required answer strings; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4 | 60.8 |
| 2 | DeepSeek V3.2 | 56.3 |
| 3 | GLM-5 | 50.2 |
| 4 | Gemini 2.5 Pro | 48 |
| 5 | GPT-4o | 41.6 |
| 6 | Qwen 3 235B A22B | 41.2 |
| 7 | GPT-OSS-120B | 34.1 |
| 8 | Llama 4 Scout | 26.9 |
| 9 | Gemma 3 27B | 15.8 |
Interactive version: theaggregate.ai/benchmark?slug=geonatureagent-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-29.