CSR-Bench (Hallucination) - Original: leaderboard

Metric: Alignment accuracy (%) on the original question about the image; CSR-Bench's hallucination subset: 1,411 images, each asked an original image-grounded question and three misleading variants, scored by a GPT-5-nano judge (1 when the response gives the image-grounded target or rejects the miscategorized request), temperature 0, 8,192-token cap; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro62.4#77
2Qwen 3 VL 8B (Thinking)53.8
3InternVL3-14B52.6#494
4Qwen 3 VL 8B Instruct51.9#401
5Qwen 3 VL 4B Instruct51.9#506
6Qwen 3 VL 4B (Thinking)51.7#471 (Qwen 3 VL 4B)
7Gemini 2.5 Flash50.1#237
8InternVL3-8B48.5#606
9Qwen 2.5 VL 7B Instruct46.4#643
10Seed-1.645.8#257
11Gemma 3 12B (IT)40.3#655
12Llama 3.2 11B Instruct38.4#1112
13Gemma 3 4B (IT)32.7#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=csr-bench-hallucination-original · How It Works · Data refreshed daily, snapshot 2026-10-11.