NonoBench — leaderboard

Nonogram puzzle solving benchmark testing LLM spatial reasoning and constraint satisfaction through Japanese logic puzzles.

Metric: Overall Accuracy (%). Source: www.nonobench.com. Status: saturation imminent. 48 models tracked.

Top models

#ModelScore
1GPT-5.4 (xHigh)63.3
2Gemini 3.1 Pro (Preview) (High)63.3
3GPT-5.4 (High)60
4Claude Opus 4.5 (High)56.7
5GPT-5.2 (High)53.3
6Gemini 3 Pro (Preview) (High)53.3
7GPT-5.2 (xHigh)50
8GPT-5.4 (Low)46.7
9GPT-5.2 (Low)46.7
10Gemini 3 Flash (Preview) (High)46.7
11Gemini 3.1 Pro (Preview) (Low)40
12GPT-OSS-120B (High)36.7
13DeepSeek V3.2 Speciale36.7
14DeepSeek V3.2 (High)36.7
15Grok 433.3

Interactive version: theaggregate.ai/benchmark?slug=nonobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.