CPGBench - Adherence: leaderboard

Metric: Adherence rate (%): share of conversations in which the conversation is cut just before the clinician applies the recommendation and the model continues it and its reply applies the recommendation, on CPGBench's 32,155 recommendations extracted from 3,418 clinical practice guidelines (nine regions, 24 specialties), each embedded in a synthetic multi-turn clinician-patient conversation, scored by a GPT-4o judge against the source recommendation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-563.18#91
2Qwen 3 32B45.89#424
3GPT-4o41.75#333
4Qwen 3 4B 2507 Instruct37.3#745
5Llama 3 8B Instruct21.77#1115

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cpgbench-adherence · How It Works · Data refreshed daily, snapshot 2026-10-11.