MultiGlobeQA (Gold Triples with Code Agent): leaderboard
Metric: Exact-match accuracy (%; oracle Tier 3: the gold triples given to the Python CodeAgent with retrieval tools disabled; exact match with answer-type tolerances (geodesic or relative tolerance, set equality, polygon IoU of at least 0.5) over the 4,979 true-premise questions of the English text small split, refusals and non-answers counted incorrect; greedy decoding, output constrained to the answer-type JSON schema; mean of three seeds). Source: arxiv.org. Saturation forecast: Around April 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 27B FP8 | 61.6 |
| 2 | Gemini 3 Flash (Preview) | 60.2 |
| 3 | Gemma 3 27B (IT) | 45.3 |
Interactive version: theaggregate.ai/benchmark?slug=multiglobeqa-gold-triples-with-code-agent · How It Works · Data refreshed daily, snapshot 2026-09-26.