MAG (Set-of-Mark) - Gated Guide Score: leaderboard

Metric: Gated guide score (%): zero on a failed task and 0.4 + 0.6 times the guide-overlap score on a solved one, averaged over the 171 test tasks with reference guides, Set-of-Mark element selection grounding; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.

Top models

#ModelScore
1GPT-5.5 (Low)20.6
2Gemini 3.5 Flash (Low)20
3Claude Sonnet 4.6 (Low)15.4

Interactive version: theaggregate.ai/benchmark?slug=mag-set-of-mark-gated-guide-score · How It Works · Data refreshed daily, snapshot 2026-09-29.