MI-CXR: leaderboard

Metric: Overall accuracy over the TEL, ICR and GTS task families, five-option multiple-choice accuracy (%, times 100; chance 20) on MI-CXR, longitudinal reasoning over five-visit chest X-ray timelines from MIMIC-CXR (5,311 questions), zero-shot single-step prompting with all five images, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.

Top models

#ModelScore
1GPT-5.241.1
2Gemini 3 Pro (Preview)38.7
3Claude Sonnet 4.531.5
4MedGemma-27B-IT29.9
5Lingshu-32B24.7
6MedGemma-4B23.7

Interactive version: theaggregate.ai/benchmark?slug=mi-cxr · How It Works · Data refreshed daily, snapshot 2026-10-07.