GACL - Battleship — leaderboard

Game Agent Coding League: Battleship sub-benchmark. LLMs write code agents to play Battleship against each other.

Metric: Normalized Score (0-100). Source: gameagentcodingleague.com. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini85.94
2GPT-5.585.38
3GPT-5.485
4Kimi K2.678.47
5GLM-5.171.19
6DeepSeek V4 Pro70.94
7Claude Opus 4.764.03
8Gemini 3.1 Pro (Preview)61.62
9DeepSeek V4 Flash61.41
10MiMo-V2.5-Pro59.44
11Claude Sonnet 4.659.25
12Claude Opus 4.658.81
13GPT-OSS-120B54.97
14Gemini 3 Flash (Preview)49.5
15Qwen 3.6 Max Preview41.44

Interactive version: theaggregate.ai/benchmark?slug=gacl-battleship · How the rankings work · Data refreshed daily, snapshot 2026-07-22.