SurgiQ - Robotic Surgery: leaderboard

Metric: Robotic surgery questions (1,643): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B69.08
2Qwen 2.5 72B Instruct68.17
3Qwen 2.5 32B Instruct65.73
4Qwen 2.5 32B65.67
5Ministral-3-14B-Reasoning-251264.09
6GPT-OSS-120B62.75
7DeepSeek R1 Distill Qwen 32B61.84
8GPT-OSS-20B57.15
9Ministral-3-8B-Instruct-251256.79
10DeepSeek R1 Distill Qwen 14B56.66
11K2 Think V253.68
12BioMistral-7B48.87
13Gemma 2 9B (IT)46.93
14MedGemma-4B-IT45.71
15Gemma 2 2B (IT)37.01

Interactive version: theaggregate.ai/benchmark?slug=surgiq-robotic-surgery · How It Works · Data refreshed daily, snapshot 2026-09-29.