Opus Magnum Bench — leaderboard

Puzzle-solving benchmark based on Opus Magnum campaign levels, scoring models by whether they synthesize valid alchemy-machine solutions and how close those solutions are to human-best efficiency.

Metric: Human-normalized score (%). Source: opusmagnumbench.com. Status: saturation imminent. 22 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)70.86
2Claude Fable 5 (High)60.15
3GPT-5.5 (xHigh)53.56
4GPT-5.6 Terra (xHigh)47.72
5GPT-5.5 (Medium)37.79
6Gemini 3.1 Pro (Preview) (High)31.11
7Claude Sonnet 5 (High)26.55
8GLM-5.2 (High)25.39
9Gemini 3.5 Flash (High)24.49
10Claude Opus 4.8 (High)22.81
11Grok 4.5 (High)18.19
12GPT-5.6 Luna (xHigh)17.8
13Gemini 3 Flash (High)15.39
14DeepSeek V4 Pro13.39
15Claude Opus 4.812.66

Interactive version: theaggregate.ai/benchmark?slug=opus-magnum-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.