CompliBench - Healthcare: leaderboard

Metric: Conversation-level accuracy (%): share of healthcare conversations in which the judge predicts both the governing guideline and the violation label correctly at every assistant turn, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro57.11
2DeepSeek V3.2 (Thinking)39.68
3Claude Sonnet 4.639.41
4Qwen 3.5 Plus36.47
5GLM-533.26
6GPT-528.44
7Qwen 3 Max25.69
8Qwen 3 30B A3B25
9Kimi K2.522.94
10GPT-4o17.66
11Qwen 3 32B8.49
12Qwen 3 14B5.5
13Qwen 3 8B1.83
14Qwen 3 4B1.83
15GPT-4o Mini0.23

Interactive version: theaggregate.ai/benchmark?slug=complibench-healthcare · How It Works · Data refreshed daily, snapshot 2026-10-07.