MM-JudgeBias - Image Misalignment: leaderboard

Metric: Bias-Deviation (0 to 1) on MM-JudgeBias (about 200 image-query-response triplets per bias type, 1,804 in all, each judged on a 1-10 scale before and after a controlled perturbation; mean of three runs): mean drop of the judge's score, normalized by the largest possible drop, when the image is replaced with an unrelated one; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 30 models tracked.

Top models

#ModelScore
1Claude Opus 4.50.97
2Claude Opus 4.5 (Thinking)0.96
3Gemini 3 Pro (High)0.96
4Claude Sonnet 4.5 (Thinking)0.91
5Gemini 2.5 Pro0.9
6Claude Sonnet 4.50.86
7Claude Haiku 4.50.82
8Claude Haiku 4.5 (Thinking)0.81
9Gemini 2.5 Flash0.76
10Qwen 3 VL 30B A3B Instruct0.65
11Gemini 2.5 Flash Lite (Thinking)0.63
12Qwen 3 VL 8B Instruct0.62
13Gemini 2.5 Flash Lite0.59
14Qwen 3 VL 8B (Thinking)0.53
15Qwen 3 VL 30B A3B (Thinking)0.48

Interactive version: theaggregate.ai/benchmark?slug=mm-judgebias-image-misalignment · How It Works · Data refreshed daily, snapshot 2026-10-07.