ShatterMed-QA (English, Easy): leaderboard

Metric: Accuracy (%) on the 5,923 English Easy multiple-choice clinical vignettes synthesized from a hub-pruned (k-shattered) medical knowledge graph, with the bridging entity masked and a hard distractor derived from its sibling node; zero-shot, closed-book; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.1 Mini98.31#346
2Grok 4 Fast (Reasoning)98.18#242 (Grok 4 Fast)
3Grok 4.1 Fast (Reasoning)98.08#208 (Grok 4.1 Fast)
4GPT-5 Mini97.45#176
5Grok 4 Fast (Non-reasoning)96.96#242 (Grok 4 Fast)
6GPT-4.1 Nano96.25#716
7GPT-5 Nano91.59#415
8Qwen3-14B-Base90.04#610
9Falcon3-10B-Base85.8#1191
10InternLM3-8B-Instruct80.2#847
11Llama 3.1 8B79.4#1139
12Yi-1.5-9B79.23#1230
13c4ai-command-r7B-12-202477.7#1145
14Gemma 2 9B77.17#940
15Grok 4.1 Fast (Non-reasoning)74.35#208 (Grok 4.1 Fast)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=shattermed-qa-english-easy · How It Works · Data refreshed daily, snapshot 2026-10-11.