ClinicalMC (English): leaderboard

Metric: Mean of 11 task scores (%; triage accuracy, examination recall, diagnosis F1, LLM-judged diagnosis basis and differential diagnosis, treatment and per-course IoU; GPT-4o-mini patient and examiner agents, gold inputs from earlier stages, mean of three runs; 5,804 English cases from PMC-Patients). Source: arxiv.org. Saturation forecast: Around 2032. 21 models tracked.

Top models

#ModelScore
1Mixtral 8x22B Instruct (v0.1)49.8
2Qwen Turbo49.68
3GPT-5 Mini46.88
4GPT-4o Mini46.23
5Qwen 2.5 32B Instruct45.98
6Qwen 3 Next 80B A3B Instruct45.59
7Qwen 2.5 14B Instruct45.38
8Qwen 2.5 72B Instruct45.15
9Mistral 7B Instruct (v0.3)45.06
10Qwen 2.5 7B Instruct44.98
11Llama 3.3 70B Instruct44.52
12Llama 3.2 3B Instruct36.47

Interactive version: theaggregate.ai/benchmark?slug=clinicalmc-english · How It Works · Data refreshed daily, snapshot 2026-09-26.