FREAK - Attribute: leaderboard

Metric: Accuracy (%) on the attribute questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 18 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro56.12#145
2O3 (High)49.89#121 (O3)
3InternVL3-78B48.49#345
4GPT-4.148.24#240
5Gemini 2.5 Flash48.02#237
6GLM-4.5V47.89#339
7O4 Mini47.22#172
8Qwen 2.5 VL 72B Instruct46.58#364
9InternVL3-38B46.46#395
10Qwen 2.5 VL 32B Instruct42.63#443
11Phi-4 Multimodal Instruct36.6#896
12Claude Sonnet 4 (Thinking)33.95#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-attribute · How It Works · Data refreshed daily, snapshot 2026-10-11.