GeoAgentBench (Base Agent): leaderboard

Metric: Parameter Execution Accuracy (%): share of ground-truth steps whose last matching tool call has semantically equivalent parameters, on 53 spatial analysis tasks over 117 atomic GIS tools in an executable sandbox (30 steps per task), agents under the base agent paradigm (zero-shot tool scheduling without execution feedback); higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash43.02
2DeepSeek V340.41
3Claude Sonnet 4.639.84
4GPT-4o Mini35.39
5GPT-4o32.54
6Llama 3.1 8B Instruct24.83
7Qwen 2.5 7B Instruct12.42

Interactive version: theaggregate.ai/benchmark?slug=geoagentbench-base-agent · How It Works · Data refreshed daily, snapshot 2026-10-07.