MapSatisfyBench - Factual Faithfulness: leaderboard

Metric: Factual-faithfulness rate (%): share of required factual claims in the final response supported by the tool outputs; underspecified map-service requests rebuilt from real anonymized user behavior chains, answered by a ReAct agent with 22 deterministic replayed map tools and a simulated user; GPT-5.3 simulates the user and judges (majority of three judgments), temperature 1; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Thinking)95.81
2GPT-4.192.15
3Qwen 3 30B A3B (Non-reasoning)92.14
4Claude Sonnet 4.691.93
5Qwen 3 8B (Non-reasoning)91.83
6Qwen 3.6 Plus (Thinking)90.84
7Claude Opus 4.690.77
8Qwen 3.6 Plus (Non-reasoning)89.85
9DeepSeek V4 Pro (Non-reasoning)89.33
10DeepSeek V3.2 (Non-reasoning)88.25
11Qwen 3 235B A22B (Non-reasoning)83.77

Interactive version: theaggregate.ai/benchmark?slug=mapsatisfybench-factual-faithfulness · How It Works · Data refreshed daily, snapshot 2026-09-29.