ReactBench (Hallucination): leaderboard

Metric: React-Score (%): question-type-weighted accuracy (short answer 0.5, multiple choice and yes/no 0.25 each), averaged over standard and chain-of-thought prompting (thinking-only models: chain-of-thought) and weighted over the four tasks (0.3, 0.2, 0.2, 0.3), hallucination-inducing questions (multiple choice, short answer, yes/no) about images edited by Qwen-Image-Edit-2511; an incorrect answer counts as a hallucination; temperature 0.2, at most 1,024 tokens, LLM answer matching; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 13 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B (Thinking)65.3
2Qwen 3 VL 32B Instruct61.9
3Qwen 3.5 27B59.5
4InternVL3-8B56.5
5Qwen 2.5 VL 72B Instruct54.2
6Qwen 3.5 9B52.4
7Qwen 2.5 VL 7B Instruct49.5

Interactive version: theaggregate.ai/benchmark?slug=reactbench-hallucination · How It Works · Data refreshed daily, snapshot 2026-10-07.