ADRD-Bench - Unified QA: leaderboard

Metric: Exact-match accuracy (0-1, shown times 100) on the 1,438 Alzheimer's-disease and dementia-related items consolidated from MedMCQA, MedQA, HEAD-QA, MEDEC, MedBullets, MedHallu and PubMedQA (multiple-choice and medical error detection), pooled over all items; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 36 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.1 70B Instruct92.98#548
2GPT-5.5 (2026-04-23)92.77#21
3Claude Opus 4.792.21#45
4Gemini 3.1 Pro (Preview)91.66#54
5Claude Opus 4.5 (20251101)91.1#70
6GPT-5.289.43#105
7Qwen 3 Max (2025-09-23)89.08#249
8Qwen 3 235B A22B 2507 Instruct87.69#291
9Llama 3 8B Instruct87.27#1115
10Llama 3.1 8B Instruct83.31#1018
11Qwen 2.5 72B Instruct83.03#436
12Qwen 3 30B A3B 2507 Instruct82.68#464
13Grok 4.1 Fast (Non-reasoning)82.06#208 (Grok 4.1 Fast)
14Qwen 2.5 14B Instruct75.17#634
15Phi-3-medium-4k-instruct74.76#939

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=adrd-bench-unified-qa · How It Works · Data refreshed daily, snapshot 2026-10-11.