ReactBench (Hallucination) - Counterfactual Attribute (Standard): leaderboard

Metric: Accuracy (%) on Counterfactual Attribute (7,668 questions on 639 images with an edited attribute) with standard prompting, hallucination-inducing questions (multiple choice, short answer, yes/no) about images edited by Qwen-Image-Edit-2511; an incorrect answer counts as a hallucination; temperature 0.2, at most 1,024 tokens, LLM answer matching; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 11 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B56.3
2Qwen 3 VL 32B Instruct55.4
3InternVL3-8B54.4
4Qwen 2.5 VL 72B Instruct53.4
5Qwen 2.5 VL 7B Instruct49.2
6Qwen 3.5 9B48.7

Interactive version: theaggregate.ai/benchmark?slug=reactbench-hallucination-counterfactual-attribute-standard · How It Works · Data refreshed daily, snapshot 2026-10-07.