Opus Magnum Bench: leaderboard

Puzzle-solving benchmark based on Opus Magnum campaign levels, scoring models by whether they synthesize valid alchemy-machine solutions and how close those solutions are to human-best efficiency.

Metric: Human-normalized score (%). Source: opusmagnumbench.com. Status: years away from saturation. 22 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)70.86
2Claude Fable 5 (High)60.15
3GPT-5.5 (xHigh)53.56
4Claude Fable 5 (Low)49.28
5GPT-5.6 Terra (xHigh)47.72
6GPT-5.5 (Medium)37.79
7Gemini 3.1 Pro (Preview) (High)31.11
8Claude Sonnet 5 (High)26.55
9GLM-5.2 (High)25.39
10Gemini 3.5 Flash (High)24.49
11Claude Opus 4.8 (High)22.81
12Grok 4.5 (High)18.19
13GPT-5.6 Luna (xHigh)17.8
14Gemini 3 Flash (High)15.39
15DeepSeek V4 Pro13.39

Interactive version: theaggregate.ai/benchmark?slug=opus-magnum-bench · How It Works · Data refreshed daily, snapshot 2026-09-05.