GameLogicBench - Repo: leaderboard
Metric: Solve rate (%; the 23 Repo tasks (a mechanic inside a real open-source project); one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude-Opus-5 (Claude Code, high) | 43.48 |
| 2 | GPT-5.6-Sol (Claude Code, high) | 34.78 |
| 3 | DeepSeek-V4-Pro-0813 (Claude Code, high) | 34.78 |
| 4 | GLM-5.2 (Claude Code, high) | 30.43 |
| 5 | DeepSeek-V4-Flash-0731 (Claude Code, high) | 26.09 |
| 6 | GPT-5.6-Sol (Codex CLI, high) | 26.09 |
| 7 | GPT-5.6-Sol (OpenCode, high) | 26.09 |
| 8 | Qwen-3.8-Max (Claude Code, high) | 21.74 |
| 9 | GPT-5.6-Terra (Codex CLI, high) | 21.74 |
| 10 | GLM-5.2 (OpenCode, high) | 21.74 |
Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench-repo · How It Works · Data refreshed daily, snapshot 2026-09-26.