MultiGlobeQA (Reasoning): leaderboard
Metric: Exact-match accuracy (%; Tier 2: the same prompt with explicit reasoning before answering (native reasoning for Gemini-3-Flash and Qwen3.5-35B, prompted chain of thought for Qwen3.5-27B and Gemma-3-27B); exact match with answer-type tolerances (geodesic or relative tolerance, set equality, polygon IoU of at least 0.5) over the 4,979 true-premise questions of the English text small split, refusals and non-answers counted incorrect; greedy decoding, output constrained to the answer-type JSON schema; mean of three seeds). Source: arxiv.org. Saturation forecast: Around November 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 24.6 |
| 2 | Gemma 3 27B (IT) | 3.8 |
| 3 | Qwen 3.5 27B FP8 | 3.8 |
Interactive version: theaggregate.ai/benchmark?slug=multiglobeqa-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.