ShatterMed-QA (English, Hard): leaderboard

Metric: Accuracy (%) on the 1,692 English Hard multiple-choice clinical vignettes synthesized from a hub-pruned (k-shattered) medical knowledge graph, with the bridging entity masked and a hard distractor derived from its sibling node; zero-shot, closed-book; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1Grok 4 Fast (Reasoning)97.87#242 (Grok 4 Fast)
2Grok 4.1 Fast (Reasoning)97.58#208 (Grok 4.1 Fast)
3GPT-4.1 Mini97.28#346
4GPT-5 Mini96.81#176
5Grok 4.1 Fast (Non-reasoning)96.1#208 (Grok 4.1 Fast)
6Grok 4 Fast (Non-reasoning)96.04#242 (Grok 4 Fast)
7GPT-4.1 Nano93.85#716
8GPT-5 Nano89.83#415
9Qwen3-14B-Base86.47#610
10Falcon3-10B-Base83.7#1191
11InternLM3-8B-Instruct78.44#847
12Llama 3.1 8B72.3#1139
13c4ai-command-r7B-12-202471.59#1145
14Yi-1.5-9B69.17#1230
15Gemma 2 9B55.82#940

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=shattermed-qa-english-hard · How It Works · Data refreshed daily, snapshot 2026-10-11.