CompliBench - Insurance - Guideline Accuracy: leaderboard

Metric: Strict guideline accuracy (%): share of compliant assistant turns in the insurance domain for which the judge predicts the correct governing guideline and no violation, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-591.91
2Gemini 3 Pro91.68
3Qwen 3.5 Plus90.8
4Qwen 3 Max89.82
5Claude Sonnet 4.685.71
6DeepSeek V3.2 (Thinking)85.5
7GLM-584.62
8Qwen 3 30B A3B82.87
9GPT-4o82.7
10Kimi K2.578.67
11Qwen 3 32B65.05
12GPT-4o Mini56.76
13Qwen 3 4B51.01
14Qwen 3 8B48.63
15Qwen 3 14B41.97

Interactive version: theaggregate.ai/benchmark?slug=complibench-insurance-guideline-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.