HoloCount - Relative Canonical Orientation: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting objects inside a region named by a canonical direction such as the bottom-left quadrant (84 questions); higher is better). Source: arxiv.org. Saturation forecast: Around August 2027. 30 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B82.1
2Qwen 3.5 27B82.1
3Gemini 3.1 Pro (Preview)81
4Qwen 3.5 35B A3B78.6
5Gemini 3 Flash (Preview)78.6
6Qwen 3.5 122B A10B77.4
7Kimi K2.576.2
8Qwen 3.5 9B76.2
9Kimi K2.676.2
10GPT-5.573.8
11Qwen 3 VL 32B Instruct72.6
12Qwen 3.5 4B71.4
13Qwen 2.5 VL 32B Instruct71.4
14Claude Opus 4.769
15Claude Sonnet 4.667.9

Interactive version: theaggregate.ai/benchmark?slug=holocount-relative-canonical-orientation · How It Works · Data refreshed daily, snapshot 2026-09-29.