SAGE (Service Agent): leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, averaged over the six scenarios, over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.572.62
2Claude Sonnet 4.571.56
3Gemini 3 Pro (Preview)71.4
4DeepSeek V3.271.29
5Kimi K2.568.71
6DeepSeek R168.39
7MiniMax-M2.168.33
8Gemini 3 Flash (Preview)68.28
9GLM-4.767.25
10Gemini 2.5 Pro66.9
11GPT-4.166.79
12DeepSeek V366.54
13Seed 1.866.14
14Qwen 2.5 32B Instruct65.34
15Qwen 2.5 72B Instruct65.27

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent · How It Works · Data refreshed daily, snapshot 2026-10-07.