CSR-Bench (Hallucination) - Wrong Premise: leaderboard

Metric: Alignment accuracy (%) when a false description that contradicts the image is prepended to the question; CSR-Bench's hallucination subset: 1,411 images, each asked an original image-grounded question and three misleading variants, scored by a GPT-5-nano judge (1 when the response gives the image-grounded target or rejects the miscategorized request), temperature 0, 8,192-token cap; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro58.7#77
2Qwen 3 VL 8B (Thinking)51.3
3InternVL3-14B49.8#494
4Qwen 3 VL 4B (Thinking)48.3#471 (Qwen 3 VL 4B)
5Qwen 3 VL 8B Instruct48.1#401
6Gemini 2.5 Flash46.2#237
7InternVL3-8B45.3#606
8Qwen 3 VL 4B Instruct44.3#506
9Seed-1.642.1#257
10Qwen 2.5 VL 7B Instruct40#643
11Gemma 3 12B (IT)36.2#655
12Llama 3.2 11B Instruct35.2#1112
13Gemma 3 4B (IT)33.4#971

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=csr-bench-hallucination-wrong-premise · How It Works · Data refreshed daily, snapshot 2026-10-11.