BunchCount - Group Counting: leaderboard

Metric: Mean absolute error (semantic groups per image, lower is better; group-level test split of 310 real images with 14 categories unseen in training, e.g. bunches of grapes, stacks of plates, pairs of shoes; vision-language models zero-shot with reasoning disabled, temperature 0, one-integer answer). Source: arxiv.org. Saturation forecast: Around August 2027. 8 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Non-reasoning)5.46
2Qwen 3.7 Max (Non-reasoning)5.55
3Qwen 3.5 27B (Non-reasoning)5.84
4Qwen 3 VL 8B Instruct10.67

Interactive version: theaggregate.ai/benchmark?slug=bunchcount-group-counting · How It Works · Data refreshed daily, snapshot 2026-09-26.