SAGE (Service Agent) - Logical Compliance: leaderboard

Metric: Logical Compliance Score (0-100): 0.4 classification accuracy + 0.4 path correctness + 0.2 action correctness against the graph, averaged over the six scenarios, over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.572.1
2DeepSeek V3.272.08
3Gemini 3 Pro (Preview)71.27
4Claude Sonnet 4.570.1
5DeepSeek R170
6Kimi K2.569.16
7GLM-4.768.68
8GPT-4.168.31
9Gemini 3 Flash (Preview)68.16
10DeepSeek V367.9
11MiniMax-M2.166.59
12Qwen 2.5 32B Instruct66.49
13Llama 3.3 70B Instruct66.06
14Qwen 3 32B65.97
15Seed 1.865.78

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-logical-compliance · How It Works · Data refreshed daily, snapshot 2026-10-07.