FREAK - OCR: leaderboard

Metric: Accuracy (%) on the ocr questions of FREAK's 1,799 questions (1,000 multiple-choice with cyclic option permutations and 799 free-form judged by GPT-5-mini) on photorealistic images edited to contradict commonsense; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 18 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro56.9#145
2GLM-4.5V56.53#339
3Gemini 2.5 Flash48.51#237
4O3 (High)47.3#121 (O3)
5InternVL3-38B46.81#395
6O4 Mini46.55#172
7Qwen 2.5 VL 72B Instruct45.7#364
8InternVL3-78B45.15#345
9Qwen 2.5 VL 32B Instruct40.74#443
10GPT-4.138.82#240
11Phi-4 Multimodal Instruct37.34#896
12Claude Sonnet 4 (Thinking)36.2#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-ocr · How It Works · Data refreshed daily, snapshot 2026-10-11.