MAG (Set-of-Mark) - Gated Guide Score: leaderboard
Metric: Gated guide score (%): zero on a failed task and 0.4 + 0.6 times the guide-overlap score on a solved one, averaged over the 171 test tasks with reference guides, Set-of-Mark element selection grounding; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (Low) | 20.6 |
| 2 | Gemini 3.5 Flash (Low) | 20 |
| 3 | Claude Sonnet 4.6 (Low) | 15.4 |
Interactive version: theaggregate.ai/benchmark?slug=mag-set-of-mark-gated-guide-score · How It Works · Data refreshed daily, snapshot 2026-09-29.