Open Medical LLM — leaderboard
Medical QA leaderboard evaluating LLMs on 9 clinical benchmarks: MedQA (USMLE), MedMCQA, PubMedQA, and 6 MMLU medical subsets. Zero-shot accuracy.
Metric: Average Accuracy (%). Source: huggingface.co. Status: saturation imminent. 181 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Daredevil-8B-abliterated | 70.86 |
| 2 | NeuralLLaMa-3-8B-DT-v0.1 | 70.43 |
| 3 | NeuralLLaMa-3-8B-ORPO-v0.3 | 70.4 |
| 4 | Llama 3 8B Ita | 70.11 |
| 5 | ChimeraLlama-3-8B-v3 | 70.04 |
| 6 | Barcenas-Llama3-8B-ORPO | 69.93 |
| 7 | Llama 3 8B | 69.9 |
| 8 | Llama-3-SauerkrautLM-8B-Instruct | 69.76 |
| 9 | suzume-llama-3-8B-multilingual | 69.3 |
| 10 | Llama 3 8B Instruct | 68.99 |
| 11 | Yi-1.5-9B | 68.89 |
| 12 | Meta-Llama-3-8B-Instruct-abliterated-v3 | 68.84 |
| 13 | Llama-3-8B-Instruct-abliterated-v2 | 68.05 |
| 14 | Yi-1.5-9B-32K | 67.73 |
| 15 | Liberated-Qwen1.5-14B | 67.66 |
Interactive version: theaggregate.ai/benchmark?slug=open-medical-llm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.