HoloCount - Linguistic Prior Conflict: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; images that contradict a strong language prior, such as a lobster with other than two claws (163 questions); higher is better). Source: arxiv.org. Saturation forecast: Around November 2027. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)73.6
2Gemini 3 Flash (Preview)69.9
3Claude Opus 4.852.1
4Qwen 3.5 27B52.1
5Qwen 3.5 397B A17B50.3
6Claude Opus 4.749.1
7GPT-5.548.5
8Kimi K2.648.5
9Qwen 3.5 122B A10B47.9
10GPT-5.446.6
11GPT-4.146.6
12Qwen 3.5 35B A3B46.6
13Kimi K2.544.8
14Gemini 2.5 Pro44.2
15Qwen 3.5 4B44.2

Interactive version: theaggregate.ai/benchmark?slug=holocount-linguistic-prior-conflict · How It Works · Data refreshed daily, snapshot 2026-09-29.