HoloCount - Visual-Prompt Region Grounding: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting only inside a region marked on the image by a visual prompt such as a box (102 questions); higher is better). Source: arxiv.org. Saturation forecast: Around July 2028. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)86.3
2Gemini 3 Flash (Preview)77.5
3Qwen 3.5 397B A17B76.5
4Qwen 3.5 122B A10B75.5
5Qwen 3.5 9B74.5
6Gemini 2.5 Pro71.6
7Qwen 3.5 35B A3B71.6
8Qwen 3.5 27B71.6
9Kimi K2.667.6
10Qwen 3.5 4B66.7
11Qwen 3 VL 32B Instruct65.7
12Qwen 2.5 VL 72B Instruct63.7
13GPT-4.161.8
14GPT-5.559.8
15Qwen 2.5 VL 32B Instruct58.8

Interactive version: theaggregate.ai/benchmark?slug=holocount-visual-prompt-region-grounding · How It Works · Data refreshed daily, snapshot 2026-09-29.