GroundedSurg - BBox IoU: leaderboard

Metric: Bounding-box IoU (0-1, times 100) between the model's predicted box and the ground-truth box, averaged over instances, on GroundedSurg (about 1,071 language-referenced surgical instrument instances in 612 images from four procedures), zero-shot: the model predicts a bounding box and centre point for the referred instrument and a frozen SAM3 backend turns them into a mask; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 17 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 2.5 VL 7B Instruct24#643
2GPT-5.214#105
3Gemma 3 27B12#596
4Qwen 3 VL 8B Instruct11#401
5GPT-4o Mini9#588
6Gemma 3 12B8#666
7Ministral 3 8B7#676
8MedGemma-4B-IT7#842

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=groundedsurg-bbox-iou · How It Works · Data refreshed daily, snapshot 2026-10-11.