SAGE (Service Agent) - Airline Refund: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Airline Refund scenario (airline refunds under rigid time-sensitive cancellation-fee policies), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.570.12
2Claude Sonnet 4.566.52
3MiniMax-M2.166.37
4DeepSeek V3.265.84
5Kimi K2.564.45
6Gemini 3 Pro (Preview)64.45
7DeepSeek R164.03
8Qwen 3 14B63.78
9GPT-4.163.68
10Qwen 2.5 14B Instruct62.96
11Llama 3.3 70B Instruct62.93
12Gemini 3 Flash (Preview)62.4
13Qwen 3 235B A22B62.37
14Seed 1.862.19
15Qwen 3 32B61.66

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-airline-refund · How It Works · Data refreshed daily, snapshot 2026-10-07.