MMRareBench - Cross-Image Evidence Alignment: leaderboard

Metric: Cross-Image Evidence Alignment track (T3) score: per-image findings, their alignment across two or more imaging modalities and the relation type, graded 0-100 by a Qwen3-VL-235B judge with a strict track-specific rubric that awards credit only when a criterion is fully met (five binary dimensions averaged); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.

Top models

#ModelScore
1GPT-570.8
2Gemini 2.5 Pro66.9
3Gemini 3 Flash (Preview)62.9
4Qwen 3 VL 235B A22B Instruct56.6
5Gemini 2.5 Flash54.9
6Claude Sonnet 4.554.7
7Claude Haiku 4.541.9
8Qwen 3 VL 30B A3B Instruct36.9
9GLM-4.6V28.7
10MedGemma-27B-IT27.3
11Qwen 3 VL 8B Instruct25.7
12Qwen 2.5 VL 32B Instruct21.2
13Qwen 2.5 VL 72B Instruct18.9
14GPT-4o15.6
15Lingshu-32B8.7

Interactive version: theaggregate.ai/benchmark?slug=mmrarebench-cross-image-evidence-alignment · How It Works · Data refreshed daily, snapshot 2026-10-07.