FREAK - Counting: leaderboard

Metric: Accuracy (%) on the counting questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 18 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro23.98#145
2O4 Mini21.96#172
3O3 (High)21.14#121 (O3)
4Gemini 2.5 Flash20.81#237
5InternVL3-78B20.12#345
6GPT-4.119.45#240
7GLM-4.5V19.41#339
8Phi-4 Multimodal Instruct18.89#896
9InternVL3-38B17.84#395
10Claude Sonnet 4 (Thinking)17.48#194 (Claude Sonnet 4)
11Qwen 2.5 VL 72B Instruct16.84#364
12Qwen 2.5 VL 32B Instruct16.74#443

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-counting · How It Works · Data refreshed daily, snapshot 2026-10-11.