SAGE (Service Agent) - Telecom Package: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Telecom Package scenario (telecom package billing and plan upgrades (linear SOP)), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 27 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.575.03
2Claude Opus 4.573.12
3DeepSeek V371.6
4DeepSeek V3.270.61
5DeepSeek R170.39
6MiniMax-M2.170.12
7Kimi K2.569.41
8Qwen 2.5 72B Instruct69.16
9Gemini 3 Pro (Preview)69.09
10Qwen 2.5 32B Instruct68.27
11Qwen 3 14B68.01
12GPT-4.166.9
13Gemini 2.5 Pro66.42
14GLM-4.766.02
15Seed 1.865.89

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-telecom-package · How It Works · Data refreshed daily, snapshot 2026-10-07.