MedCTA — leaderboard

Metric: Outcome Accuracy (self-reported). Source: benchmarklist.com. 18 models tracked.

Top models

#ModelScore
1GPT-5.431.54
2Claude Opus 4.631.32
3GPT-5.4 Mini28.31
4Qwen 3 8B27.8
5Gemini 3 Flash (Preview)25.87
6Claude Sonnet 4.625.33
7Claude Haiku 4.523.08
8Qwen 3.5 9B21.64
9GPT-5.4 Nano20.3
10Llama 3.1 8B Instruct18.94
11Llama 3.2 3B Instruct11.29
12deepseek-llm-7B-chat11
13Phi-410.65
14Mistral 7B9.4
15GPT-OSS-20B3.18

Interactive version: theaggregate.ai/benchmark?slug=medcta · How the rankings work · Data refreshed daily, snapshot 2026-07-22.