GameLogicBench: leaderboard
Metric: Solve rate (%; all 72 tasks; one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude-Opus-5 (Claude Code, high) | 52.78 |
| 2 | Qwen-3.8-Max (Claude Code, high) | 44.44 |
| 3 | GPT-5.6-Sol (Claude Code, high) | 41.67 |
| 4 | DeepSeek-V4-Pro-0813 (Claude Code, high) | 41.67 |
| 5 | GLM-5.2 (Claude Code, high) | 37.5 |
| 6 | Kimi-K3 (Claude Code, high) | 36.11 |
| 7 | GPT-5.6-Sol (Codex CLI, high) | 34.72 |
| 8 | GPT-5.6-Sol (OpenCode, high) | 34.72 |
| 9 | DeepSeek-V4-Flash-0731 (Claude Code, high) | 31.94 |
| 10 | GLM-5.2 (OpenCode, high) | 31.94 |
Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench · How It Works · Data refreshed daily, snapshot 2026-09-26.