MultiGlobeQA (Reasoning): leaderboard

Metric: Exact-match accuracy (%; Tier 2: the same prompt with explicit reasoning before answering (native reasoning for Gemini-3-Flash and Qwen3.5-35B, prompted chain of thought for Qwen3.5-27B and Gemma-3-27B); exact match with answer-type tolerances (geodesic or relative tolerance, set equality, polygon IoU of at least 0.5) over the 4,979 true-premise questions of the English text small split, refusals and non-answers counted incorrect; greedy decoding, output constrained to the answer-type JSON schema; mean of three seeds). Source: arxiv.org. Saturation forecast: Around November 2027. 4 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)24.6
2Gemma 3 27B (IT)3.8
3Qwen 3.5 27B FP83.8

Interactive version: theaggregate.ai/benchmark?slug=multiglobeqa-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.