LoMeVQA - Progress Report Generation: leaderboard
Metric: F1-RadGraph clinical efficacy (%; progress report written from an image sequence, clinical entities and relations matched against the reference; LoMeVQA-test, 2,500 expert-reviewed longitudinal chest radiograph samples over five tasks, zero-shot; an input longer than the model window prints N/A and is left out; higher is better). Source: arxiv.org. Saturation forecast: Around 2033. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Lingshu-32B | 16.45 |
| 2 | Gemini 2.5 Flash | 14.28 |
| 3 | Lingshu-7B | 13.12 |
| 4 | GPT-5 | 11.77 |
| 5 | GPT-4o | 10.79 |
| 6 | MedGemma-4B | 10.73 |
| 7 | InternVL3.5-8B | 8.43 |
| 8 | MedGemma-27B-IT | 7.34 |
| 9 | Qwen 3 VL 32B | 7.02 |
Interactive version: theaggregate.ai/benchmark?slug=lomevqa-progress-report-generation · How It Works · Data refreshed daily, snapshot 2026-09-29.