SurgiQ - Case-Based: leaderboard

Metric: Case-based questions (6,156): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B67.9
2Qwen 2.5 72B Instruct67.38
3Qwen 2.5 32B65.98
4GPT-OSS-120B64.75
5Qwen 2.5 32B Instruct64.23
6Ministral-3-14B-Reasoning-251263.22
7DeepSeek R1 Distill Qwen 32B59.9
8GPT-OSS-20B58.85
9K2 Think V258.2
10Ministral-3-8B-Instruct-251255.75
11DeepSeek R1 Distill Qwen 14B55.27
12BioMistral-7B47.53
13MedGemma-4B-IT45.26
14Gemma 2 9B (IT)43.6
15MedGemma-27B-IT39.45

Interactive version: theaggregate.ai/benchmark?slug=surgiq-case-based · How It Works · Data refreshed daily, snapshot 2026-09-29.