LMGame-Bench Tetris — leaderboard
LLM game-playing benchmark: Tetris testing spatial reasoning and piece placement optimization.
Metric: Score. Source: huggingface.co. Status: saturation imminent. 25 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4 | 125.7 |
| 2 | GPT-5 (High) | 84.3 |
| 3 | O3 (2025-04-16) | 42 |
| 4 | O1 (2024-12-17) | 35 |
| 5 | DeepSeek R1 0528 | 33.7 |
| 6 | O4 Mini (2025-04-16) | 25.3 |
| 7 | Gemini 2.5 Pro (Preview 05-06) | 23.3 |
| 8 | Grok 3 Mini Beta | 21.3 |
| 9 | Claude Opus 4 (20250514) | 20 |
| 10 | GLM-4.5 | 19.7 |
| 11 | Claude Sonnet 4 (20250514) | 19.3 |
| 12 | Kimi K2 (0711) | 17 |
| 13 | Claude 3.7 Sonnet (20250219) | 16.3 |
| 14 | Gemini 2.5 Flash (Preview 04-17) | 16.3 |
| 15 | Claude 3.5 Sonnet (20241022) | 14.7 |
Interactive version: theaggregate.ai/benchmark?slug=lmgame-bench-tetris · How the rankings work · Data refreshed daily, snapshot 2026-07-22.