FREAK - Analysis: leaderboard

Metric: Accuracy (%) on the analysis questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 18 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro35.67#145
2GPT-4.133.26#240
3Gemini 2.5 Flash32.1#237
4O3 (High)31.54#121 (O3)
5O4 Mini30.43#172
6InternVL3-78B29.65#345
7InternVL3-38B28.09#395
8Qwen 2.5 VL 72B Instruct28.09#364
9GLM-4.5V26.99#339
10Qwen 2.5 VL 32B Instruct25.8#443
11Phi-4 Multimodal Instruct25.52#896
12Claude Sonnet 4 (Thinking)24.96#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-analysis · How It Works · Data refreshed daily, snapshot 2026-10-11.