CSR-Bench (Hallucination): leaderboard

Metric: Alignment accuracy (%) averaged over the four question forms (original, wrong premise, distractor text, modality conflict); CSR-Bench's hallucination subset: 1,411 images, each asked an original image-grounded question and three misleading variants, scored by a GPT-5-nano judge (1 when the response gives the image-grounded target or rejects the miscategorized request), temperature 0, 8,192-token cap; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro51.3#77
2Qwen 3 VL 8B Instruct41.4#401
3Qwen 3 VL 8B (Thinking)40.2
4InternVL3-14B39.8#494
5Qwen 3 VL 4B (Thinking)38.4#471 (Qwen 3 VL 4B)
6Qwen 3 VL 4B Instruct37.6#506
7Gemini 2.5 Flash37.2#237
8InternVL3-8B35.7#606
9Qwen 2.5 VL 7B Instruct34.3#643
10Seed-1.632.9#257
11Gemma 3 12B (IT)30.1#655
12Llama 3.2 11B Instruct26.2#1112
13Gemma 3 4B (IT)23.9#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=csr-bench-hallucination · How It Works · Data refreshed daily, snapshot 2026-10-11.