MI-CXR - ICR Variant: leaderboard
Metric: Interval-wise Change Reasoning variant in which the pair of visits is specified and only the change type (new, resolved, progressed and similar) is asked, single abnormality, five-option multiple-choice accuracy (%, times 100; chance 20) on MI-CXR, longitudinal reasoning over five-visit chest X-ray timelines from MIMIC-CXR (5,311 questions), zero-shot single-step prompting with all five images, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 76.5 |
| 2 | Gemini 3 Pro (Preview) | 74.3 |
| 3 | MedGemma-27B-IT | 70.5 |
| 4 | Lingshu-32B | 70.5 |
| 5 | MedGemma-4B | 61.7 |
| 6 | Claude Sonnet 4.5 | 60.1 |
Interactive version: theaggregate.ai/benchmark?slug=mi-cxr-icr-variant · How It Works · Data refreshed daily, snapshot 2026-10-07.