GeoAgentBench (ReAct) - Map Verification: leaderboard
Metric: VLM verification score (0 to 100, mean of three GPT-4o judgements) comparing the agent's final map with the gold-trajectory map for data-spatial accuracy and cartographic style, on 53 spatial analysis tasks over 117 atomic GIS tools in an executable sandbox (30 steps per task), agents under the ReAct paradigm (thought, action and observation loop with runtime feedback); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 78.1 |
| 2 | DeepSeek V3 | 65.1 |
| 3 | Gemini 2.5 Flash | 64.3 |
| 4 | GPT-4o | 55.6 |
| 5 | GPT-4o Mini | 36.9 |
| 6 | Qwen 2.5 7B Instruct | 13.6 |
| 7 | Llama 3.1 8B Instruct | 7.7 |
Interactive version: theaggregate.ai/benchmark?slug=geoagentbench-react-map-verification · How It Works · Data refreshed daily, snapshot 2026-10-07.