ShatterMed-QA (Chinese, Easy): leaderboard

Metric: Accuracy (%) on the 2,616 Chinese Easy multiple-choice clinical vignettes synthesized from a hub-pruned (k-shattered) medical knowledge graph, with the bridging entity masked and a hard distractor derived from its sibling node; zero-shot, closed-book; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1Grok 4 Fast (Reasoning)98.59#242 (Grok 4 Fast)
2Grok 4.1 Fast (Reasoning)98.24#208 (Grok 4.1 Fast)
3GPT-4.1 Mini98.17#346
4GPT-5 Mini97.82#176
5Grok 4 Fast (Non-reasoning)97.82#242 (Grok 4 Fast)
6Grok 4.1 Fast (Non-reasoning)97.63#208 (Grok 4.1 Fast)
7Qwen3-14B-Base95.11#610
8GPT-4.1 Nano92.35#716
9InternLM3-8B-Instruct91.66#847
10GPT-5 Nano91.13#415
11Yi-1.5-9B88.09#1230
12c4ai-command-r7B-12-202479.03#1145
13Gemma 2 9B77.79#940
14Falcon3-10B-Base76.86#1191
15Llama 3.1 8B72.69#1139

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=shattermed-qa-chinese-easy · How It Works · Data refreshed daily, snapshot 2026-10-11.