CompliBench - Airline: leaderboard

Metric: Conversation-level accuracy (%): share of airline conversations in which the judge predicts both the governing guideline and the violation label correctly at every assistant turn, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3 Pro49.09
2GPT-547.26
3DeepSeek V3.2 (Thinking)42.07
4Qwen 3.5 Plus41.16
5GLM-531.4
6Kimi K2.525.91
7Qwen 3 Max24.39
8Qwen 3 30B A3B21.95
9Claude Sonnet 4.621.18
10GPT-4o17.07
11Qwen 3 32B8.54
12Qwen 3 4B3.66
13GPT-4o Mini2.74
14Qwen 3 14B2.13
15Qwen 3 8B1.22

Interactive version: theaggregate.ai/benchmark?slug=complibench-airline · How It Works · Data refreshed daily, snapshot 2026-10-07.