MapSatisfyBench - Acceptance Rate: leaderboard

Metric: Accepted-response score (%): per instance the explicit-completion rate times the weighted implicit-factor satisfaction, averaged over instances; underspecified map-service requests rebuilt from real anonymized user behavior chains, answered by a ReAct agent with 22 deterministic replayed map tools and a simulated user; GPT-5.3 simulates the user and judges (majority of three judgments), temperature 1; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.667.49
2DeepSeek V4 Pro (Non-reasoning)66.06
3Qwen 3.6 Plus (Thinking)62.78
4Claude Sonnet 4.661.69
5DeepSeek V3.2 (Non-reasoning)61.2
6Gemini 3.1 Pro (Preview) (Thinking)60.84
7Qwen 3.6 Plus (Non-reasoning)60.82
8Qwen 3 235B A22B (Non-reasoning)57.69
9GPT-4.156.83
10Qwen 3 8B (Non-reasoning)30.93
11Qwen 3 30B A3B (Non-reasoning)28.21

Interactive version: theaggregate.ai/benchmark?slug=mapsatisfybench-acceptance-rate · How It Works · Data refreshed daily, snapshot 2026-09-29.