CompliBench - Insurance: leaderboard

Metric: Conversation-level accuracy (%): share of insurance conversations in which the judge predicts both the governing guideline and the violation label correctly at every assistant turn, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro51.3
2Qwen 3.5 Plus40.43
3DeepSeek V3.2 (Thinking)33.91
4GLM-531.09
5GPT-530
6Claude Sonnet 4.628.75
7Qwen 3 30B A3B26.52
8Qwen 3 Max25.65
9Kimi K2.517.61
10GPT-4o11.3
11Qwen 3 32B8.91
12Qwen 3 14B3.26
13GPT-4o Mini2.17
14Qwen 3 8B1.74
15Qwen 3 4B1.09

Interactive version: theaggregate.ai/benchmark?slug=complibench-insurance · How It Works · Data refreshed daily, snapshot 2026-10-07.