GBQA: leaderboard

Metric: Recall (%) in Quality Assurance Mode (the agent also reads the game's design documents and source code) on the 124 human-verified bugs implanted in GBQA's 30 games, found by the model driving the authors' ReAct QA agent (with in-session and cross-session memory) for at most 500 interaction steps per game; a GPT-5.2 critic agent matches each report to the ground-truth bugs; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 22 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)48.39
2Qwen 3.5 397B A17B (Thinking)41.13
3Claude Opus 4.637.9
4DeepSeek R137.9
5Claude Sonnet 4.5 (Thinking)37.1
6Qwen 3 235B A22B (Thinking)35.48
7O334.68
8Qwen 3 32B (Thinking)33.87
9Claude Sonnet 4.532.26
10Kimi K2.5 (Thinking)28.23
11Qwen 3 8B (Thinking)24.19
12Qwen 3.5 397B A17B (Non-reasoning)24.19
13GPT-5.2 (Non-reasoning)22.58
14Kimi K2.5 (Non-reasoning)20.97
15DeepSeek V3.2 (Non-reasoning)20.16

Interactive version: theaggregate.ai/benchmark?slug=gbqa · How It Works · Data refreshed daily, snapshot 2026-10-07.