CompliBench - Healthcare - Guideline Accuracy: leaderboard

Metric: Strict guideline accuracy (%): share of compliant assistant turns in the healthcare domain for which the judge predicts the correct governing guideline and no violation, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-592.86
2Gemini 3 Pro92.39
3Qwen 3.5 Plus92.39
4GPT-4o90.41
5DeepSeek V3.2 (Thinking)89.95
6Qwen 3 Max88.24
7Claude Sonnet 4.687.33
8GLM-586.18
9Qwen 3 30B A3B84.55
10Kimi K2.579.35
11Qwen 3 32B62.42
12Qwen 3 4B54.43
13GPT-4o Mini51.13
14Qwen 3 8B49.34
15Qwen 3 14B40.49

Interactive version: theaggregate.ai/benchmark?slug=complibench-healthcare-guideline-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.