GameLogicBench - Combo: leaderboard

Metric: Solve rate (%; the 28 Combo tasks (several interacting mechanics); one sample per model-scaffold configuration with effort=high, a 3,600 s limit and network egress sealed; a task counts as solved only if the submitted Godot 4.4 project passes every tick-level state assertion in every judge-selected scenario and seed). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 20 models tracked.

Top models

#ModelScore
1Claude-Opus-5 (Claude Code, high)53.57
2Qwen-3.8-Max (Claude Code, high)50
3Kimi-K3 (Claude Code, high)42.86
4DeepSeek-V4-Pro-0813 (Claude Code, high)39.29
5GPT-5.6-Sol (Claude Code, high)35.71
6GPT-5.6-Sol (OpenCode, high)35.71
7GLM-5.2 (Claude Code, high)32.14
8DeepSeek-V4-Flash-0731 (Claude Code, high)32.14
9GPT-5.6-Sol (Codex CLI, high)32.14
10DeepSeek-V4-Pro (Claude Code, high)28.57

Interactive version: theaggregate.ai/benchmark?slug=gamelogicbench-combo · How It Works · Data refreshed daily, snapshot 2026-09-26.