ShatterMed-QA (Chinese, Hard): leaderboard

Metric: Accuracy (%) on the 327 Chinese Hard multiple-choice clinical vignettes synthesized from a hub-pruned (k-shattered) medical knowledge graph, with the bridging entity masked and a hard distractor derived from its sibling node; zero-shot, closed-book; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.1 Mini96.94#346
2Grok 4.1 Fast (Reasoning)96.64#208 (Grok 4.1 Fast)
3Grok 4 Fast (Reasoning)96.64#242 (Grok 4 Fast)
4GPT-5 Mini96.33#176
5Grok 4.1 Fast (Non-reasoning)96.33#208 (Grok 4.1 Fast)
6Grok 4 Fast (Non-reasoning)96.02#242 (Grok 4 Fast)
7Qwen3-14B-Base92.35#610
8InternLM3-8B-Instruct89.3#847
9GPT-5 Nano86.54#415
10GPT-4.1 Nano84.71#716
11Yi-1.5-9B80.73#1230
12Gemma 2 9B79.51#940
13c4ai-command-r7B-12-202474.31#1145
14Llama 3.1 8B67.89#1139
15Falcon3-10B-Base63.61#1191

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=shattermed-qa-chinese-hard · How It Works · Data refreshed daily, snapshot 2026-10-11.