GameLogicBench - Atom: leaderboard
Metric: Solve rate (%; the 21 Atom tasks (one mechanic in a minimal game); one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude-Opus-5 (Claude Code, high) | 61.9 |
| 2 | Qwen-3.8-Max (Claude Code, high) | 61.9 |
| 3 | GPT-5.6-Sol (Claude Code, high) | 57.14 |
| 4 | Qwen-3.8-Max (Codex CLI, high) | 57.14 |
| 5 | GLM-5.2 (OpenCode, high) | 57.14 |
| 6 | DeepSeek-V4-Pro-0813 (Claude Code, high) | 52.38 |
| 7 | GLM-5.2 (Claude Code, high) | 52.38 |
| 8 | Kimi-K3 (Claude Code, high) | 52.38 |
| 9 | Kimi-K3 (Codex CLI, high) | 52.38 |
| 10 | Qwen-3.8-Max (OpenCode, high) | 52.38 |
Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench-atom · How It Works · Data refreshed daily, snapshot 2026-09-26.