GTO Wizard Benchmark: leaderboard

Metric: AIVAT luck-adjusted win rate in big blinds per hundred hands against GTO Wizard AI over 5,000 heads-up no-limit hands (open scale, 0 is break-even against the bot); zero-shot text prompts, no tools; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 15 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.4 (xHigh)-17.84#76 (GPT-5.4)
2Claude Opus 4.6-20.35#60
3Claude Opus 4.5-22.26#79
4Gemini 3 Pro-30.06#77
5Gemini 3.1 Pro (Preview)-30.83#54
6Gemini 2.5 Pro-39.19#145
7Kimi K2.5-41.43#139
8GPT-5.4 (Non-reasoning)-57.5#76 (GPT-5.4)
9GPT-4o-60.9#333
10GPT-5.4 Mini (Non-reasoning)-107.87#202 (GPT-5.4 Mini)
11GPT-4-136.17#423
12GPT-5.4 Nano-189.73#320

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=gto-wizard-benchmark · How It Works · Data refreshed daily, snapshot 2026-10-11.