MM-JudgeBias - Visual Transformation: leaderboard

Metric: Bias-Conformity (0 to 1) on MM-JudgeBias (about 200 image-query-response triplets per bias type, 1,804 in all, each judged on a 1-10 scale before and after a controlled perturbation; mean of three runs): mean score invariance (1 minus the absolute score change over its largest possible value) when the image receives meaning-preserving augmentations; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 30 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (High)0.89
2GPT-4.1 Mini0.89
3GPT-5 Mini (High)0.88
4Qwen 3 VL 8B Instruct0.87
5GPT-5.1 (High)0.87
6Qwen 3 VL 8B (Thinking)0.87
7Gemini 2.5 Flash0.86
8Gemini 2.0 Flash Lite0.86
9Gemini 2.5 Pro0.86
10Qwen 2.5 VL 72B Instruct0.86
11Qwen 3 VL 30B A3B Instruct0.85
12O4 Mini (High)0.85
13Qwen 3 VL 30B A3B (Thinking)0.85
14Claude Opus 4.5 (Thinking)0.84
15Claude Sonnet 4.5 (Thinking)0.83

Interactive version: theaggregate.ai/benchmark?slug=mm-judgebias-visual-transformation · How It Works · Data refreshed daily, snapshot 2026-10-07.