MapSatisfyBench - Tool Selection: leaderboard

Metric: Tool-selection score (%): set overlap between the valid tools the agent called and the annotated required tools, penalizing missing and redundant calls; underspecified map-service requests rebuilt from real anonymized user behavior chains, answered by a ReAct agent with 22 deterministic replayed map tools and a simulated user; GPT-5.3 simulates the user and judges (majority of three judgments), temperature 1; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.649.42
2Claude Sonnet 4.649.09
3GPT-4.146.96
4Gemini 3.1 Pro (Preview) (Thinking)45.6
5Qwen 3.6 Plus (Thinking)45.36
6Qwen 3.6 Plus (Non-reasoning)44.93
7DeepSeek V4 Pro (Non-reasoning)43.82
8DeepSeek V3.2 (Non-reasoning)43.79
9Qwen 3 235B A22B (Non-reasoning)42.36
10Qwen 3 8B (Non-reasoning)33.49
11Qwen 3 30B A3B (Non-reasoning)28.59

Interactive version: theaggregate.ai/benchmark?slug=mapsatisfybench-tool-selection · How It Works · Data refreshed daily, snapshot 2026-09-29.