FREAK - Detection: leaderboard

Metric: Accuracy (%) on the detection questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 18 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.150.25#240
2O3 (High)48.96#121 (O3)
3Gemini 2.5 Pro47.85#145
4GLM-4.5V47.85#339
5Qwen 2.5 VL 72B Instruct47.23#364
6O4 Mini44.51#172
7Gemini 2.5 Flash44.04#237
8InternVL3-78B43.03#345
9InternVL3-38B40.06#395
10Phi-4 Multimodal Instruct39.49#896
11Qwen 2.5 VL 32B Instruct38.33#443
12Claude Sonnet 4 (Thinking)29.93#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-detection · How It Works · Data refreshed daily, snapshot 2026-10-11.