LoMeVQA - Progress Description: leaderboard
Metric: F1-RadGraph clinical efficacy (%; free-text description of the change between two radiographs, clinical entities and relations matched against the reference; LoMeVQA-test, 2,500 expert-reviewed longitudinal chest radiograph samples over five tasks, zero-shot; higher is better). Source: arxiv.org. Saturation forecast: Around 2038. 16 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Lingshu-32B | 26.67 |
| 2 | Gemini 2.5 Flash | 23.55 |
| 3 | MedGemma-4B | 20.95 |
| 4 | Lingshu-7B | 19.78 |
| 5 | InternVL3.5-8B | 19.4 |
| 6 | GPT-5 | 18.3 |
| 7 | GPT-4o | 14.36 |
| 8 | MedGemma-27B-IT | 12.93 |
| 9 | Qwen 3 VL 32B | 8.41 |
Interactive version: theaggregate.ai/benchmark?slug=lomevqa-progress-description · How It Works · Data refreshed daily, snapshot 2026-09-29.