MMRareBench - Examination Suggestion: leaderboard

Metric: Examination Suggestion track (T4) score: prioritized next-step investigations for a case narrative truncated before the gold tests, graded 0-100 by a Qwen3-VL-235B judge with a strict track-specific rubric that awards credit only when a criterion is fully met (eight 0-2 dimensions normalized by 16); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 23 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.585.9
2Gemini 2.5 Pro85.5
3Qwen 3 VL 235B A22B Instruct83.2
4Gemini 3 Flash (Preview)81.8
5Claude Haiku 4.580.6
6GPT-577.9
7Gemini 2.5 Flash66.9
8Qwen 3 VL 30B A3B Instruct63.1
9MedGemma-27B-IT56
10Qwen 3 VL 8B Instruct43.9
11GPT-4o42.8
12Qwen 2.5 VL 72B Instruct42.3
13Qwen 2.5 VL 32B Instruct42
14Lingshu-32B36.9
15MedGemma-4B-IT33.1

Interactive version: theaggregate.ai/benchmark?slug=mmrarebench-examination-suggestion · How It Works · Data refreshed daily, snapshot 2026-10-07.