MM-JudgeBias - Image Dominance: leaderboard

Metric: Bias-Deviation (0 to 1) on MM-JudgeBias (about 200 image-query-response triplets per bias type, 1,804 in all, each judged on a 1-10 scale before and after a controlled perturbation; mean of three runs): mean drop of the judge's score, normalized by the largest possible drop, when the query is replaced with empty text; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 30 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (Thinking)0.59
2Claude Opus 4.50.57
3Claude Sonnet 4.5 (Thinking)0.54
4Gemini 2.5 Pro0.54
5Claude Haiku 4.5 (Thinking)0.53
6O4 Mini (High)0.45
7Claude Sonnet 4.50.39
8Claude Haiku 4.50.37
9Gemini 2.5 Flash Lite0.37
10Gemini 2.0 Flash Lite0.36
11O3 (High)0.35
12Qwen 3 VL 30B A3B (Thinking)0.34
13Gemini 2.5 Flash Lite (Thinking)0.34
14Gemini 2.5 Flash0.29
15Gemini 3 Pro (High)0.28

Interactive version: theaggregate.ai/benchmark?slug=mm-judgebias-image-dominance · How It Works · Data refreshed daily, snapshot 2026-10-07.