CompliBench - Insurance - Violation Detection: leaderboard

Metric: Violation detection accuracy (%): share of violating assistant turns in the insurance domain that the judge flags as violations, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro95.26
2Kimi K2.593.09
3GLM-591.85
4Claude Sonnet 4.688.86
5DeepSeek V3.2 (Thinking)86.26
6Qwen 3.5 Plus86.26
7GPT-583.93
8Qwen 3 Max78.96
9Qwen 3 30B A3B78.11
10Qwen 3 32B57.84
11GPT-4o54.89
12Qwen 3 14B43.71
13Qwen 3 4B36.88
14Qwen 3 8B31.99
15GPT-4o Mini26.94

Interactive version: theaggregate.ai/benchmark?slug=complibench-insurance-violation-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.