GroundedSurg: leaderboard

Metric: Mask IoU (0-1, times 100) between the projected mask and the ground-truth instrument mask, averaged over instances, on GroundedSurg (about 1,071 language-referenced surgical instrument instances in 612 images from four procedures), zero-shot: the model predicts a bounding box and centre point for the referred instrument and a frozen SAM3 backend turns them into a mask; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 17 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 2.5 VL 7B Instruct20#643
2GPT-5.29#105
3Qwen 3 VL 8B Instruct9#401
4Gemma 3 12B7#666
5Gemma 3 27B6#596
6GPT-4o Mini4#588
7Ministral 3 8B3#676
8MedGemma-4B-IT2#842

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=groundedsurg · How It Works · Data refreshed daily, snapshot 2026-10-11.