GameLogicBench - Repo: leaderboard

Metric: Solve rate (%; the 23 Repo tasks (a mechanic inside a real open-source project); one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.

Top models

#ModelScore
1Claude-Opus-5 (Claude Code, high)43.48
2GPT-5.6-Sol (Claude Code, high)34.78
3DeepSeek-V4-Pro-0813 (Claude Code, high)34.78
4GLM-5.2 (Claude Code, high)30.43
5DeepSeek-V4-Flash-0731 (Claude Code, high)26.09
6GPT-5.6-Sol (Codex CLI, high)26.09
7GPT-5.6-Sol (OpenCode, high)26.09
8Qwen-3.8-Max (Claude Code, high)21.74
9GPT-5.6-Terra (Codex CLI, high)21.74
10GLM-5.2 (OpenCode, high)21.74

Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench-repo · How It Works · Data refreshed daily, snapshot 2026-09-26.