GameLogicBench - Atom: leaderboard

Metric: Solve rate (%; the 21 Atom tasks (one mechanic in a minimal game); one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.

Top models

#ModelScore
1Claude-Opus-5 (Claude Code, high)61.9
2Qwen-3.8-Max (Claude Code, high)61.9
3GPT-5.6-Sol (Claude Code, high)57.14
4Qwen-3.8-Max (Codex CLI, high)57.14
5GLM-5.2 (OpenCode, high)57.14
6DeepSeek-V4-Pro-0813 (Claude Code, high)52.38
7GLM-5.2 (Claude Code, high)52.38
8Kimi-K3 (Claude Code, high)52.38
9Kimi-K3 (Codex CLI, high)52.38
10Qwen-3.8-Max (OpenCode, high)52.38

Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench-atom · How It Works · Data refreshed daily, snapshot 2026-09-26.