MedSPOT: leaderboard
Metric: Task completion accuracy (%): share of MedSPOT's clinical GUI tasks (2-3 interdependent click steps each, in ten DICOM viewers, segmentation tools and web viewers) in which every step's predicted click lands in the target box, evaluated strictly in sequence with termination at the first miss; one predicted point per step except GUI-Actor, scored with its top-5 candidates and a 14-pixel tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 3 VL 8B Instruct | 35.05 | #401 |
| 2 | Qwen 2.5 VL 7B Instruct | 12.62 | #643 |
| 3 | GPT-5 | 2.8 | #91 |
| 4 | Gemma 3 27B (IT) | 0 | #509 |
| 5 | GPT-4o Mini | 0 | #588 |
| 6 | Llama 3.2 11B Instruct | 0 | #1112 |
| 7 | Qwen 2 VL 7B Instruct | 0 | #816 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=medspot · How It Works · Data refreshed daily, snapshot 2026-10-11.