GeoAgentBench (ReAct) - Tool Retrieval F1: leaderboard

Metric: Tools-Any-Order F1 (%) between the tools the agent invokes and the ground-truth tool set, on 53 spatial analysis tasks over 117 atomic GIS tools in an executable sandbox (30 steps per task), agents under the ReAct paradigm (thought, action and observation loop with runtime feedback); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.683.59
2GPT-4o80.65
3Gemini 2.5 Flash78.55
4DeepSeek V378.16
5GPT-4o Mini65.71
6Llama 3.1 8B Instruct47.64
7Qwen 2.5 7B Instruct28.52

Interactive version: theaggregate.ai/benchmark?slug=geoagentbench-react-tool-retrieval-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.