SAGE (Service Agent) - Ecommerce Refund: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Ecommerce Refund scenario (ecommerce refunds (deep decision tree on product status and credit level)), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.572.73
2Qwen 2.5 32B Instruct72.07
3DeepSeek V3.271.51
4Claude Sonnet 4.571.21
5Gemini 3 Pro (Preview)70.99
6GPT-4.170.32
7Seed 1.870.25
8DeepSeek R170.16
9Kimi K2.569.69
10DeepSeek V369.62
11Gemini 2.5 Pro69.46
12Qwen 2.5 7B Instruct69.05
13Qwen 2.5 3B Instruct68.92
14Llama 3.3 70B Instruct68.6
15MiniMax-M2.168.33

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-ecommerce-refund · How It Works · Data refreshed daily, snapshot 2026-10-07.