CritPt — leaderboard

Research-level physics reasoning benchmark with composite challenges designed by active physics researchers.

Metric: Accuracy (self-reported). Source: benchmarklist.com. Status: saturation imminent. 328 models tracked.

Top models

#ModelScore
1GPT-5.5 Pro30.6
2GPT-5.4 Pro (xHigh)30
3Claude Mythos 528.6
4Claude Fable 5 (Max)28.6
5GPT-5.527.1
6Gemini 3 Deep Think25.7
7GPT-5.423.4
8Claude Opus 4.8 (Max)20.9
9Gemini 3.1 Pro (Preview)17.7
10GPT-5.3 Codex16.9
11GLM-5.216.7
12Qwen 3.7 Max (Max)13.4
13Gemini 3.5 Flash13.1
14DeepSeek V4 Pro (Max)12.9
15Claude Opus 4.6 (Max)12.6

Interactive version: theaggregate.ai/benchmark?slug=critpt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.