CompliBench - Airline - Violation Detection: leaderboard

Metric: Violation detection accuracy (%): share of violating assistant turns in the airline domain that the judge flags as violations, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro90.61
2Kimi K2.586.88
3GLM-586.46
4DeepSeek V3.2 (Thinking)84.81
5Qwen 3.5 Plus82.32
6Claude Sonnet 4.681.76
7GPT-581.08
8Qwen 3 Max71.41
9Qwen 3 30B A3B69.75
10GPT-4o50.55
11Qwen 3 32B41.44
12Qwen 3 14B32.04
13Qwen 3 8B28.18
14Qwen 3 4B24.86
15GPT-4o Mini19.61

Interactive version: theaggregate.ai/benchmark?slug=complibench-airline-violation-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.