SurgiQ - Neurosurgery: leaderboard

Metric: Neurosurgery questions (2,204): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B68.51
2Qwen 2.5 72B Instruct67.97
3Qwen 2.5 32B66.33
4Ministral-3-14B-Reasoning-251266.29
5GPT-OSS-120B65.74
6Qwen 2.5 32B Instruct64.88
7DeepSeek R1 Distill Qwen 32B60.53
8GPT-OSS-20B58.98
9K2 Think V257.08
10DeepSeek R1 Distill Qwen 14B56.99
11Ministral-3-8B-Instruct-251254.08
12BioMistral-7B50.64
13MedGemma-4B-IT46.87
14Gemma 2 9B (IT)44.92
15MedGemma-27B-IT39.11

Interactive version: theaggregate.ai/benchmark?slug=surgiq-neurosurgery · How It Works · Data refreshed daily, snapshot 2026-09-29.