KataGo-Bench-1K — leaderboard
Go next-move prediction benchmark with 1,000 19x19 board positions sampled from KataGo annotation data. Models read the move history and rendered board state, then output the next move; scoring is accuracy against KataGo candidate moves.
Metric: Next-move Accuracy (%). Source: arxiv.org. Status: saturation imminent. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.7 Sonnet | 34.3 |
| 2 | O1 Mini | 27.3 |
| 3 | DeepSeek R1 | 17.6 |
| 4 | Qwen 2.5 7B Instruct | 8 |
| 5 | Qwen 2.5 32B Instruct | 6.8 |
| 6 | DeepSeek R1 Distill Qwen 32B | 4.7 |
| 7 | Qwen 2.5 32B | 1.5 |
| 8 | Qwen 2.5 7B | 1.4 |
| 9 | DeepSeek-R1-Distill-Qwen-7B | 0.6 |
Interactive version: theaggregate.ai/benchmark?slug=katago-bench-1k · How the rankings work · Data refreshed daily, snapshot 2026-07-22.