gg-bench — leaderboard

Zero-sum grid-based game competition benchmark from UC Berkeley. Evaluates domain-general capabilities and strategic reasoning without relying on human-curated QA tasks.

Metric: Games Won. Source: github.com. Status: years away from saturation. 7 models tracked.

Top models

#ModelScore
1O144.27
2O3 Mini30.27
3DeepSeek R122.27
4Llama 3.3 70B17.77
5Claude 3.7 Sonnet7.07
6GPT-4o Mini2.77
7GPT-4o2.57

Interactive version: theaggregate.ai/benchmark?slug=gg-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.