CPGBench - Content Detection: leaderboard
Metric: Content detection rate (%): share of conversations in which the model reads the whole conversation and must report any clinical recommendation it contains and names the embedded one, on CPGBench's 32,155 recommendations extracted from 3,418 clinical practice guidelines (nine regions, 24 specialties), each embedded in a synthetic multi-turn clinician-patient conversation, scored by a GPT-4o judge against the source recommendation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Qwen 3 32B | 84.77 | #424 |
| 2 | GPT-4o | 82.65 | #333 |
| 3 | GPT-5 | 79.47 | #91 |
| 4 | Llama 3 8B Instruct | 79.38 | #1115 |
| 5 | Qwen 3 4B 2507 Instruct | 72.01 | #745 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=cpgbench-content-detection · How It Works · Data refreshed daily, snapshot 2026-10-11.