HoloCount - Null-Target Prompting: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; questions about objects absent from the image, where the correct count is zero (250 questions); higher is better). Source: arxiv.org. Saturation forecast: Estimated already saturated. 30 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct98
2Claude Sonnet 4.597.2
3Qwen 3 VL 8B Instruct96.4
4Claude Opus 4.795.6
5Kimi K2.695.6
6Qwen 2.5 VL 7B Instruct95.6
7Claude Sonnet 4.695.2
8Qwen 2.5 VL 72B Instruct95.2
9Claude Opus 4.694.4
10Kimi K2.594.4
11Qwen 2.5 VL 32B Instruct93.6
12Gemini 2.5 Pro93.4
13InternVL3.5-8B92.8
14Qwen 3.5 397B A17B90.8
15Qwen 3.5 122B A10B90.8

Interactive version: theaggregate.ai/benchmark?slug=holocount-null-target-prompting · How It Works · Data refreshed daily, snapshot 2026-09-29.