MapSatisfyBench - Implicit Intent Satisfaction: leaderboard

Metric: Implicit-intent satisfaction rate (%): evidence-weighted satisfaction of the unstated decision factors (hard constraints binary, soft preferences graded); underspecified map-service requests rebuilt from real anonymized user behavior chains, answered by a ReAct agent with 22 deterministic replayed map tools and a simulated user; GPT-5.3 simulates the user and judges (majority of three judgments), temperature 1; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 15 models tracked.

Top models

#ModelScore
1Claude Opus 4.671.7
2DeepSeek V4 Pro (Non-reasoning)70.17
3Qwen 3.6 Plus (Thinking)67.83
4Claude Sonnet 4.667.25
5DeepSeek V3.2 (Non-reasoning)66.1
6Qwen 3.6 Plus (Non-reasoning)65.85
7Gemini 3.1 Pro (Preview) (Thinking)64.69
8Qwen 3 235B A22B (Non-reasoning)62.37
9GPT-4.161.04
10Qwen 3 8B (Non-reasoning)37.16
11Qwen 3 30B A3B (Non-reasoning)36.35

Interactive version: theaggregate.ai/benchmark?slug=mapsatisfybench-implicit-intent-satisfaction · How It Works · Data refreshed daily, snapshot 2026-09-29.