SAGE (Service Agent) - Logistics Delivery: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Logistics Delivery scenario (logistics exceptions such as lost or delayed parcels and insurance claims), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScore
1Claude Opus 4.571.25
2DeepSeek V3.270.87
3MiniMax-M2.168.54
4Gemini 3 Pro (Preview)68.05
5Claude Sonnet 4.567.3
6Gemini 3 Flash (Preview)67.1
7GLM-4.766.5
8DeepSeek R165.75
9Kimi K2.565.74
10Gemini 2.5 Pro64.92
11Qwen 3 235B A22B64.77
12Seed 1.864.01
13Qwen 2.5 72B Instruct63.16
14Qwen 2.5 32B Instruct62.99
15GPT-4.162.69

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-logistics-delivery · How It Works · Data refreshed daily, snapshot 2026-10-07.