Chess APO Bench: leaderboard

Metric: Puzzle accuracy (%) on the 559-puzzle Lichess test split with the unoptimized base prompt (every position must match the single reference move), mean of three runs. Source: arxiv.org. Saturation forecast: Around February 2027. 9 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash (Medium)55.4
2GPT-5.6 Luna (Low)31.72
3DeepSeek V4 Pro (0813)14.49
4Claude Haiku 4.511.63
5Qwen 3.8 27B8.47
6GPT-4o Mini5.84

Interactive version: theaggregate.ai/benchmark?slug=chess-apo-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.