gg-bench — leaderboard
Zero-sum grid-based game competition benchmark from UC Berkeley. Evaluates domain-general capabilities and strategic reasoning without relying on human-curated QA tasks.
Metric: Games Won. Source: github.com. Status: years away from saturation. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O1 | 44.27 |
| 2 | O3 Mini | 30.27 |
| 3 | DeepSeek R1 | 22.27 |
| 4 | Llama 3.3 70B | 17.77 |
| 5 | Claude 3.7 Sonnet | 7.07 |
| 6 | GPT-4o Mini | 2.77 |
| 7 | GPT-4o | 2.57 |
Interactive version: theaggregate.ai/benchmark?slug=gg-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.