GACL - Wizard — leaderboard

Game Agent Coding League: Wizard card game sub-benchmark. LLMs write code agents to play the trick-taking card game Wizard.

Metric: Normalized Score (0-100). Source: gameagentcodingleague.com. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1GPT-5.571.22
2GPT-5.466.07
3Gemini 3.1 Pro (Preview)61.77
4DeepSeek V4 Pro59.7
5Claude Opus 4.758.11
6Kimi K2.657.41
7Claude Opus 4.656.04
8Gemini 3 Flash (Preview)55.22
9Qwen 3.6 Plus55.15
10Qwen 3.6 Max Preview54.73
11MiMo-V2.5-Pro51.98
12GPT-OSS-120B48.69
13Gemma 4 31B (IT)47.92
14Claude Sonnet 4.647.61
15DeepSeek V4 Flash46.98

Interactive version: theaggregate.ai/benchmark?slug=gacl-wizard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.