ZebraLogic — leaderboard

Logic grid puzzle benchmark from Allen AI testing systematic constraint satisfaction and deductive reasoning across varying difficulty levels.

Metric: Puzzle Accuracy (%). Source: huggingface.co. Status: saturation imminent. 62 models tracked.

Top models

#ModelScore
1O3 Mini (2025-01-31) (High)91.7
2O3 Mini (Medium)88.9
3O1 (2024-12-17)81
4DeepSeek R178.7
5O3 Mini (Low)74.8
6O1 Preview (2024-09-12)71.4
7O1 Mini (2024-09-12)59.7
8DeepSeek V342.1
9Claude 3.5 Sonnet (20241022)36.2
10Claude 3.5 Sonnet (20240620)33.4
11GPT-4o (2024-08-06)31.7
12Gemini 1.5 Pro (Preview 0827)30.5
13Mistral Large 2 (Jul)29
14GPT-4 Turbo28.4
15GPT-4o (2024-05-13)28.2

Interactive version: theaggregate.ai/benchmark?slug=zebralogic · How the rankings work · Data refreshed daily, snapshot 2026-07-22.