MageBench Season 1 — leaderboard
MageBench Season 1 evaluates model capability on games & game agents tasks from the linked upstream source with Rating as the primary reported metric.
Metric: Rating (self-reported). Source: benchmarklist.com. Status: saturation imminent. 35 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Medium) | 1747 |
| 2 | GPT-5.2 (Medium) | 1737 |
| 3 | DeepSeek V3.2 | 1682 |
| 4 | GPT-5.4 (Medium) | 1658 |
| 5 | Gemini 3 Flash (Preview) (Medium) | 1622 |
| 6 | O3 (Medium) | 1609 |
| 7 | Qwen 3 235B A22B | 1594 |
| 8 | Llama 4 Maverick | 1590 |
| 9 | Gemini 3.1 Flash Lite | 1578 |
| 10 | GPT-5.2 | 1547 |
| 11 | GPT-4o Mini | 1546 |
| 12 | GPT-5 (Medium) | 1536 |
| 13 | GPT-5 Mini (Medium) | 1516 |
| 14 | Mistral Large | 1501 |
| 15 | GPT-5 Nano (Low) | 1499 |
Interactive version: theaggregate.ai/benchmark?slug=magebench-season-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.