CSR-Bench (Hallucination) - Distractor Text: leaderboard

Metric: Alignment accuracy (%) when fabricated common-knowledge text that points away from the visual evidence is added; CSR-Bench's hallucination subset: 1,411 images, each asked an original image-grounded question and three misleading variants, scored by a GPT-5-nano judge (1 when the response gives the image-grounded target or rejects the miscategorized request), temperature 0, 8,192-token cap; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro45.2#77
2Qwen 3 VL 8B (Thinking)43.6
3Qwen 3 VL 4B (Thinking)40.7#471 (Qwen 3 VL 4B)
4Qwen 3 VL 8B Instruct38.8#401
5Qwen 3 VL 4B Instruct33.9#506
6Gemma 3 12B (IT)32.3#655
7InternVL3-14B32.1#494
8Gemini 2.5 Flash30.5#237
9Qwen 2.5 VL 7B Instruct29.8#643
10InternVL3-8B28.7#606
11Seed-1.625.4#257
12Gemma 3 4B (IT)24.6#971
13Llama 3.2 11B Instruct18.5#1112

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=csr-bench-hallucination-distractor-text · How It Works · Data refreshed daily, snapshot 2026-10-11.