FREAK (Multiple Choice) - Consistency Accuracy: leaderboard

Metric: Consistency accuracy (%): share of FREAK's 1,000 multiple-choice questions answered correctly under all six cyclic permutations of the options; reasoning enabled where the model supports it; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 23 models tracked.

Top models

#ModelScoreOverall rank
1InternVL3-78B35.9#345
2InternVL3-38B35#395
3GLM-4.5V34.3#339
4Gemini 2.5 Pro33#145
5O3 (High)32.4#121 (O3)
6GPT-4.131.5#240
7Qwen 2.5 VL 72B Instruct30.4#364
8Qwen 2.5 VL 32B Instruct28.3#443
9O4 Mini27.6#172
10Qwen 2.5 VL 7B Instruct26#643
11InternVL3-14B25.3#494
12Gemini 2.5 Flash25.1#237
13InternVL3-8B24.8#606
14Phi-4 Multimodal Instruct19.2#896
15Claude Sonnet 4 (Thinking)18.5#194 (Claude Sonnet 4)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=freak-multiple-choice-consistency-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.