ReactBench (Hallucination) - Dense Counting (Standard): leaderboard

Metric: Accuracy (%) on Dense Counting (12,248 questions on 1,531 crowded images) with standard prompting, hallucination-inducing questions (multiple choice, short answer, yes/no) about images edited by Qwen-Image-Edit-2511; an incorrect answer counts as a hallucination; temperature 0.2, at most 1,024 tokens, LLM answer matching; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 11 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B69.3
2Qwen 3 VL 32B Instruct66.9
3InternVL3-8B64.2
4Qwen 2.5 VL 72B Instruct63.3
5Qwen 2.5 VL 7B Instruct58.6
6Qwen 3.5 9B57.2

Interactive version: theaggregate.ai/benchmark?slug=reactbench-hallucination-dense-counting-standard · How It Works · Data refreshed daily, snapshot 2026-10-07.