MI-CXR: leaderboard
Metric: Overall accuracy over the TEL, ICR and GTS task families, five-option multiple-choice accuracy (%, times 100; chance 20) on MI-CXR, longitudinal reasoning over five-visit chest X-ray timelines from MIMIC-CXR (5,311 questions), zero-shot single-step prompting with all five images, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 41.1 |
| 2 | Gemini 3 Pro (Preview) | 38.7 |
| 3 | Claude Sonnet 4.5 | 31.5 |
| 4 | MedGemma-27B-IT | 29.9 |
| 5 | Lingshu-32B | 24.7 |
| 6 | MedGemma-4B | 23.7 |
Interactive version: theaggregate.ai/benchmark?slug=mi-cxr · How It Works · Data refreshed daily, snapshot 2026-10-07.