GACL - Connect4 — leaderboard

Game Agent Coding League: Connect4 sub-benchmark. LLMs write code agents to play Connect Four against each other.

Metric: Normalized Score (0-100). Source: gameagentcodingleague.com. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)76.3
2Claude Opus 4.773.23
3GLM-5.171.2
4GPT-5.569.64
5GPT-5.469.32
6Kimi K2.666.46
7Claude Opus 4.665.57
8Claude Sonnet 4.657.81
9Qwen 3.6 Max Preview56.41
10Qwen 3.6 35B A3B55.83
11Qwen 3.6 Plus55
12DeepSeek V4 Pro54.37
13MiMo-V2.5-Pro45.05
14Qwen 3.6 27B43.54
15MiniMax-M2.742.76

Interactive version: theaggregate.ai/benchmark?slug=gacl-connect4 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.