MedSPOT: leaderboard

Metric: Task completion accuracy (%): share of MedSPOT's clinical GUI tasks (2-3 interdependent click steps each, in ten DICOM viewers, segmentation tools and web viewers) in which every step's predicted click lands in the target box, evaluated strictly in sequence with termination at the first miss; one predicted point per step except GUI-Actor, scored with its top-5 candidates and a 14-pixel tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 8B Instruct35.05#401
2Qwen 2.5 VL 7B Instruct12.62#643
3GPT-52.8#91
4Gemma 3 27B (IT)0#509
5GPT-4o Mini0#588
6Llama 3.2 11B Instruct0#1112
7Qwen 2 VL 7B Instruct0#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medspot · How It Works · Data refreshed daily, snapshot 2026-10-11.