SAGE (Service Agent) - Online Education: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Online Education scenario (online education contract disputes and refund-risk control), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)82.15
2DeepSeek V3.278.17
3Gemini 3 Flash (Preview)75.62
4Claude Opus 4.575.22
5Claude Sonnet 4.574.24
6MiniMax-M2.172.55
7GLM-4.772.36
8GPT-4.171.69
9Kimi K2.571.15
10Gemini 2.5 Pro70.28
11DeepSeek R169.4
12DeepSeek V368.77
13Qwen 2.5 72B Instruct68.76
14Qwen 3 235B A22B68.02
15Qwen 3 32B67.56

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-online-education · How It Works · Data refreshed daily, snapshot 2026-10-07.