ReactBench (Hallucination) - Relational Erasure (Standard): leaderboard

Metric: Accuracy (%) on Relational Erasure (19,140 questions on 1,595 images where an object or relation was edited out) with standard prompting, hallucination-inducing questions (multiple choice, short answer, yes/no) about images edited by Qwen-Image-Edit-2511; an incorrect answer counts as a hallucination; temperature 0.2, at most 1,024 tokens, LLM answer matching; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 11 models tracked.

Top models

#ModelScore
1Qwen 3 VL 32B Instruct64.5
2Qwen 2.5 VL 72B Instruct55.6
3InternVL3-8B51.6
4Qwen 3.5 27B48
5Qwen 2.5 VL 7B Instruct47
6Qwen 3.5 9B43.2

Interactive version: theaggregate.ai/benchmark?slug=reactbench-hallucination-relational-erasure-standard · How It Works · Data refreshed daily, snapshot 2026-10-07.