SurgiQ - Best-Option: leaderboard

Metric: Best-option questions (2,085), several options partly correct: accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct65.58
2Qwen 2.5 72B64.76
3GPT-OSS-120B62.7
4Qwen 2.5 32B Instruct61.88
5Qwen 2.5 32B61.74
6Ministral-3-14B-Reasoning-251259.91
7GPT-OSS-20B56.65
8DeepSeek R1 Distill Qwen 32B55.98
9K2 Think V255.11
10Ministral-3-8B-Instruct-251254.58
11DeepSeek R1 Distill Qwen 14B51.9
12BioMistral-7B47.05
13MedGemma-4B-IT47
14Gemma 2 9B (IT)41.81
15MedGemma-27B-IT36.25

Interactive version: theaggregate.ai/benchmark?slug=surgiq-best-option · How It Works · Data refreshed daily, snapshot 2026-09-29.