ClinicalMC (Chinese): leaderboard

Metric: Mean of 11 task scores (%; triage accuracy, examination recall, diagnosis F1, LLM-judged diagnosis basis and differential diagnosis, treatment and per-course IoU; GPT-4o-mini patient and examiner agents, gold inputs from earlier stages, mean of three runs; 1,275 Chinese records from MedEureka). Source: arxiv.org. Saturation forecast: Around January 2028. 22 models tracked.

Top models

#ModelScore
1Qwen 3 Next 80B A3B Instruct48.17
2GPT-5 Mini42.04
3Qwen Turbo38.62
4Qwen 2.5 32B Instruct38.36
5Qwen 2.5 7B Instruct38.27
6Llama 3.3 70B Instruct37.93
7Qwen 2.5 14B Instruct37.81
8Qwen 2.5 72B Instruct37.59
9GPT-4o Mini34.84
10Mixtral 8x22B Instruct (v0.1)30.25
11Llama 3.2 3B Instruct22.01
12Mistral 7B Instruct (v0.3)20.65

Interactive version: theaggregate.ai/benchmark?slug=clinicalmc-chinese · How It Works · Data refreshed daily, snapshot 2026-09-26.