CompliBench - Airline - Guideline Accuracy: leaderboard

Metric: Strict guideline accuracy (%): share of compliant assistant turns in the airline domain for which the judge predicts the correct governing guideline and no violation, of an LLM judge on CompliBench's synthesized multi-turn dialogues with injected, adversarially optimized guideline violations, mean of four runs at default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1GPT-591.65
2Gemini 3 Pro87.71
3Qwen 3.5 Plus87.34
4DeepSeek V3.2 (Thinking)86.54
5Qwen 3 Max85.11
6GPT-4o83.51
7GLM-579.31
8Qwen 3 30B A3B78.3
9Claude Sonnet 4.678.18
10Kimi K2.573.35
11Qwen 3 32B55.74
12GPT-4o Mini51.7
13Qwen 3 8B44.41
14Qwen 3 4B43.24
15Qwen 3 14B36.81

Interactive version: theaggregate.ai/benchmark?slug=complibench-airline-guideline-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.