CPGBench - Title Grounding: leaderboard

Metric: Title grounding rate (%): share of conversations in which the model reads the whole conversation and must report any clinical recommendation it contains and also gives the title of the guideline it comes from, on CPGBench's 32,155 recommendations extracted from 3,418 clinical practice guidelines (nine regions, 24 specialties), each embedded in a synthetic multi-turn clinician-patient conversation, scored by a GPT-4o judge against the source recommendation; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-529.68#91
2Qwen 3 32B7.96#424
3GPT-4o7.91#333
4Llama 3 8B Instruct5.33#1115
5Qwen 3 4B 2507 Instruct4.17#745

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cpgbench-title-grounding · How It Works · Data refreshed daily, snapshot 2026-10-11.