Step Game (Lechmazur) — leaderboard

Tests spatial reasoning - models must track relative positions of characters over 10 steps (e.g., 'A is to the left of B'). TrueSkill ranking from multi-game tournaments across 75+ models.

Metric: TrueSkill μ. Source: github.com. Status: saturation imminent. 75 models tracked.

Top models

#ModelScore
1GPT-5 (Medium)5.49
2O3 (Medium)5.32
3GPT-5.1 (Medium)5.29
4Gemini 3 Pro (Preview)5.03
5O14.14
6Gemini 2.5 Flash3.99
7Gemini 2.5 Pro3.82
8Grok 4.1 Fast (Reasoning)3.78
9Gemini 2.5 Pro (Preview 03-25)3.72
10DeepSeek V3.2 Exp3.67
11O3 Mini (High)3.64
12Grok 43.47
13O3 Mini (Medium)3.38
14Claude Sonnet 4.5 (Thinking 16K)3.37
15Grok 3 Mini (High)3.24

Interactive version: theaggregate.ai/benchmark?slug=step-game-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.