FREAK - Position: leaderboard

Metric: Accuracy (%) on the position questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 18 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro43.41#145
2GPT-4.139.37#240
3O3 (High)38.77#121 (O3)
4InternVL3-78B38.41#345
5Qwen 2.5 VL 72B Instruct38.22#364
6GLM-4.5V37.41#339
7InternVL3-38B37.21#395
8O4 Mini35.87#172
9Gemini 2.5 Flash35.46#237
10Phi-4 Multimodal Instruct32.77#896
11Qwen 2.5 VL 32B Instruct31.77#443
12Claude Sonnet 4 (Thinking)25.45#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-position · How It Works · Data refreshed daily, snapshot 2026-10-11.