ClockBench — leaderboard

Visual reasoning benchmark testing if AI can read analog clocks. 180 clocks, 720 questions covering time reading, arithmetic, rotation, and timezone conversion. Human baseline: 90.7%.

Metric: Accuracy (%). Source: clockbench.ai. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1Human Expert100
2Median Human90.7
3GPT-5.6 Sol (Max)66.7
4GPT-5.4 (High)50.6
5GPT-5.5 (High)46.1
6Qwen 3 VL 235B A22B Instruct39.4
7Claude Fable 535
8Gemini 3.1 Pro (Preview)32.2
9Gemini 3.5 Flash31.1
10Gemini 3 Pro28.9
11Grok 4.521.7
12Gemini 2.5 Pro18.9
13Claude Opus 4.715
14GPT-5.2 (High)15
15Qwen 3 VL 235B A22B (Thinking)14.4

Interactive version: theaggregate.ai/benchmark?slug=clockbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.