CompliBench - Healthcare - Violation Detection: leaderboard

Metric: Violation detection accuracy (%): share of violating assistant turns in the healthcare domain that the judge flags as violations, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro94.8
2Claude Sonnet 4.690.37
3Kimi K2.589.45
4GLM-588.15
5DeepSeek V3.2 (Thinking)82.49
6Qwen 3.5 Plus80.5
7Qwen 3 30B A3B77.83
8GPT-577.68
9Qwen 3 Max72.71
10GPT-4o57.03
11Qwen 3 32B43.43
12Qwen 3 14B31.35
13Qwen 3 8B25.99
14Qwen 3 4B22.02
15GPT-4o Mini15.37

Interactive version: theaggregate.ai/benchmark?slug=complibench-healthcare-violation-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.