SurgiQ - Negative: leaderboard

Metric: Negative (all-except) questions (2,153): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B68.31
2Qwen 2.5 72B Instruct68.03
3Qwen 2.5 32B Instruct67.01
4Ministral-3-14B-Reasoning-251264.96
5Qwen 2.5 32B64.82
6DeepSeek R1 Distill Qwen 32B64.36
7GPT-OSS-120B62.08
8DeepSeek R1 Distill Qwen 14B59.25
9GPT-OSS-20B48.98
10K2 Think V246.28
11Gemma 2 9B (IT)46.1
12Ministral-3-8B-Instruct-251243.12
13BioMistral-7B41.87
14MedGemma-27B-IT36.2
15DeepSeek-R1-Distill-Qwen-7B34.15

Interactive version: theaggregate.ai/benchmark?slug=surgiq-negative · How It Works · Data refreshed daily, snapshot 2026-09-29.