SnakeBench — leaderboard
LLM Battle Snake arena where AI models play the classic Snake game competitively, testing strategic planning and decision-making.
Metric: TrueSkill Rating. Source: snakebench.com. Status: saturation imminent. 318 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 38 |
| 2 | GPT-5.5 | 37.4 |
| 3 | Gemini 3.1 Pro (Preview) | 37 |
| 4 | GPT-5.1 Codex Mini | 36.5 |
| 5 | O3 | 36.4 |
| 6 | GPT-5.1 Codex Max | 36.4 |
| 7 | Claude Opus 4.8 | 35.8 |
| 8 | GPT-5 | 35 |
| 9 | Qwen 3.7 Max | 34.1 |
| 10 | O4 Mini | 33.9 |
| 11 | O4 Mini (High) | 33.7 |
| 12 | GPT-5 Codex | 33.3 |
| 13 | GPT-5.2 | 33.2 |
| 14 | Gemini 2.0 Flash | 33.2 |
| 15 | Gemini 3.5 Flash | 33.2 |
Interactive version: theaggregate.ai/benchmark?slug=snakebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.