GBQA (Player Exploring Mode): leaderboard

Metric: Recall (%) in Player Exploring Mode (the agent sees only what a player sees) on the 124 human-verified bugs implanted in GBQA's 30 games, found by the model driving the authors' ReAct QA agent (with in-session and cross-session memory) for at most 500 interaction steps per game; a GPT-5.2 critic agent matches each report to the ground-truth bugs; higher is better. Source: arxiv.org. Saturation forecast: Around September 2028. 22 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)35.48
2Claude Opus 4.631.45
3Qwen 3.5 397B A17B (Thinking)30.65
4DeepSeek R127.42
5Claude Sonnet 4.5 (Thinking)26.61
6O325
7Qwen 3 235B A22B (Thinking)25
8Qwen 3 32B (Thinking)24.19
9Claude Sonnet 4.520.97
10Kimi K2.5 (Thinking)20.16
11Qwen 3 8B (Thinking)16.13
12Qwen 3.5 397B A17B (Non-reasoning)15.32
13GPT-5.2 (Non-reasoning)14.52
14Kimi K2.5 (Non-reasoning)13.71
15DeepSeek V3.2 (Non-reasoning)12.9

Interactive version: theaggregate.ai/benchmark?slug=gbqa-player-exploring-mode · How It Works · Data refreshed daily, snapshot 2026-10-07.