SurgiQ - Critical Care and Emergency: leaderboard

Metric: Critical care and emergency surgery questions (590): accuracy (%) on SurgiQ's four-option surgical multiple-choice questions, zero-shot; the answer is the option label with the highest next-token log-probability (plain and space-prefixed labels), options shuffled once. Source: arxiv.org. Saturation forecast: Around December 2026. 35 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B68.14
2Qwen 2.5 72B Instruct67.8
3Qwen 2.5 32B64.92
4GPT-OSS-120B64.41
5Ministral-3-14B-Reasoning-251262.88
6Qwen 2.5 32B Instruct61.69
7GPT-OSS-20B60.68
8DeepSeek R1 Distill Qwen 32B59.83
9K2 Think V259.66
10Ministral-3-8B-Instruct-251252.88
11DeepSeek R1 Distill Qwen 14B51.69
12BioMistral-7B46.27
13Gemma 2 9B (IT)43.9
14MedGemma-4B-IT43.22
15MedGemma-27B-IT38.47

Interactive version: theaggregate.ai/benchmark?slug=surgiq-critical-care-and-emergency · How It Works · Data refreshed daily, snapshot 2026-09-29.