ADRD-Bench - Caregiving QA: leaderboard

Metric: Exact-match accuracy (0-1, shown times 100) on 149 caregiving questions (120 true/false and 29 multiple-choice) derived from the Aging Brain Care program's caregiver education materials and reviewed by a senior clinician, pooled over all items; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 36 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.5 (2026-04-23)95.97#21
2Claude Opus 4.795.3#45
3Gemini 3.1 Pro (Preview)93.96#54
4GPT-5.293.96#105
5Qwen 3 235B A22B 2507 Instruct93.29#291
6Llama 3.1 70B Instruct92.62#548
7Claude Opus 4.5 (20251101)92.62#70
8Qwen 3 Max (2025-09-23)91.95#249
9Gemma 3 12B (IT)90.6#655
10Qwen 2.5 72B Instruct89.93#436
11Phi-3-medium-4k-instruct89.26#939
12Grok 4.1 Fast (Non-reasoning)88.59#208 (Grok 4.1 Fast)
13Qwen 3 30B A3B 2507 Instruct87.92#464
14Qwen 3 4B 2507 Instruct87.25#745
15Qwen 2.5 7B Instruct85.91#846

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=adrd-bench-caregiving-qa · How It Works · Data refreshed daily, snapshot 2026-10-11.