ZebraLogic: leaderboard

Logic grid puzzle benchmark from Allen AI testing systematic constraint satisfaction and deductive reasoning across varying difficulty levels.

Metric: Puzzle Accuracy (%). Source: huggingface.co. Status: saturated. 62 models tracked.

Top models

#ModelScore
1O3 Mini (2025-01-31) (High)91.7
2O3 Mini (Medium)88.9
3O1 (2024-12-17)81
4DeepSeek R178.7
5O3 Mini (Low)74.8
6O1 Preview (2024-09-12)71.4
7O1 Mini (2024-09-12)59.7
8DeepSeek V342.1
9Claude 3.5 Sonnet (20241022)36.2
10Claude 3.5 Sonnet (20240620)33.4
11GPT-4o (2024-08-06)31.7
12Mistral Large 2 (Jul)29
13GPT-4 Turbo28.4
14GPT-4o (2024-05-13)28.2
15Grok 2 (1212)27.7

Interactive version: theaggregate.ai/benchmark?slug=zebralogic · How It Works · Data refreshed daily, snapshot 2026-09-05.