CPGBench - Title Grounding: leaderboard
Metric: Title grounding rate (%): share of conversations in which the model reads the whole conversation and must report any clinical recommendation it contains and also gives the title of the guideline it comes from, on CPGBench's 32,155 recommendations extracted from 3,418 clinical practice guidelines (nine regions, 24 specialties), each embedded in a synthetic multi-turn clinician-patient conversation, scored by a GPT-4o judge against the source recommendation; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5 | 29.68 | #91 |
| 2 | Qwen 3 32B | 7.96 | #424 |
| 3 | GPT-4o | 7.91 | #333 |
| 4 | Llama 3 8B Instruct | 5.33 | #1115 |
| 5 | Qwen 3 4B 2507 Instruct | 4.17 | #745 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=cpgbench-title-grounding · How It Works · Data refreshed daily, snapshot 2026-10-11.