ReactBench (Hallucination) - Relational Erasure (CoT): leaderboard

Metric: Accuracy (%) on Relational Erasure (19,140 questions on 1,595 images where an object or relation was edited out) with a step-by-step reasoning prompt, hallucination-inducing questions (multiple choice, short answer, yes/no) about images edited by Qwen-Image-Edit-2511; an incorrect answer counts as a hallucination; temperature 0.2, at most 1,024 tokens, LLM answer matching; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 13 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B (Thinking)60.7
2Qwen 3 VL 32B Instruct58.9
3Qwen 2.5 VL 7B Instruct49.3
4InternVL3-8B48.2
5Qwen 3.5 9B47.4
6Qwen 2.5 VL 72B Instruct47.2
7Qwen 3.5 27B46.4

Interactive version: theaggregate.ai/benchmark?slug=reactbench-hallucination-relational-erasure-cot · How It Works · Data refreshed daily, snapshot 2026-10-07.