CPGBench - Adherence: leaderboard
Metric: Adherence rate (%): share of conversations in which the conversation is cut just before the clinician applies the recommendation and the model continues it and its reply applies the recommendation, on CPGBench's 32,155 recommendations extracted from 3,418 clinical practice guidelines (nine regions, 24 specialties), each embedded in a synthetic multi-turn clinician-patient conversation, scored by a GPT-4o judge against the source recommendation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 63.18 | #91 |
| 2 | Qwen 3 32B | 45.89 | #424 |
| 3 | GPT-4o | 41.75 | #333 |
| 4 | Qwen 3 4B 2507 Instruct | 37.3 | #745 |
| 5 | Llama 3 8B Instruct | 21.77 | #1115 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=cpgbench-adherence · How It Works · Data refreshed daily, snapshot 2026-10-11.