FREAK (Multiple Choice): leaderboard

Metric: Accuracy (%) on FREAK's 1,000 multiple-choice questions, each asked under all six cyclic permutations of its three content options (option D fixed as none listed); reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2034. 23 models tracked.

Top models

#ModelScoreOverall rank
1O3 (High)43.98#121 (O3)
2GPT-4.143.78#240
3InternVL3-78B43.63#345
4GLM-4.5V42.78#339
5Gemini 2.5 Pro42.73#145
6InternVL3-38B42.6#395
7O4 Mini41.58#172
8InternVL3-8B41.23#606
9Gemini 2.5 Flash41#237
10Phi-4 Multimodal Instruct37.56#896
11Qwen 2.5 VL 72B Instruct36.97#364
12Qwen 2.5 VL 7B Instruct36.08#643
13InternVL3-14B36.03#494
14Qwen 2.5 VL 32B Instruct36.03#443
15Claude Sonnet 4 (Thinking)30.3#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-10-11.