SurgiQ - Reasoning: leaderboard

Metric: Reasoning questions (2,673): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct71.25
2Qwen 2.5 72B70.83
3GPT-OSS-120B67.91
4Qwen 2.5 32B67.45
5Qwen 2.5 32B Instruct67.15
6Ministral-3-14B-Reasoning-251265.69
7GPT-OSS-20B62.16
8DeepSeek R1 Distill Qwen 32B61.56
9K2 Think V260.21
10Ministral-3-8B-Instruct-251260.17
11DeepSeek R1 Distill Qwen 14B57.81
12BioMistral-7B50.53
13MedGemma-4B-IT48.61
14Gemma 2 9B (IT)44.41
15MedGemma-27B-IT40.35

Interactive version: theaggregate.ai/benchmark?slug=surgiq-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.