GeoAgentBench (Plan-and-React): leaderboard
Metric: Parameter Execution Accuracy (%): share of ground-truth steps whose last matching tool call has semantically equivalent parameters, on 53 spatial analysis tasks over 117 atomic GIS tools in an executable sandbox (30 steps per task), agents under the authors' Plan-and-React framework (global plan with step-wise reactive execution); higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3 | 47.34 |
| 2 | Claude Sonnet 4.6 | 46.11 |
| 3 | Gemini 2.5 Flash | 43.19 |
| 4 | GPT-4o | 40.69 |
| 5 | GPT-4o Mini | 36.15 |
| 6 | Llama 3.1 8B Instruct | 21.59 |
| 7 | Qwen 2.5 7B Instruct | 16.9 |
Interactive version: theaggregate.ai/benchmark?slug=geoagentbench-plan-and-react · How It Works · Data refreshed daily, snapshot 2026-10-07.