SAGE (Service Agent) - Property Service: leaderboard

Metric: Overall Assessment Score (0-100), 0.8 x logical compliance + 0.2 conversational quality, in the Property Service scenario (property service requests such as repairs and noise complaints), over SAGE customer-service dialogues: standard operating procedures formalized as dialogue graphs, an adversarial simulated user with varied intents and personas, turns 1, 5, 10, 15 and the final turn scored; three unnamed judge LLMs (majority vote and mean) plus a deterministic rule engine produce the reference; closed models through their APIs, open models served with vLLM; 0-100 scale; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 27 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.575.07
2Gemini 3 Pro (Preview)73.65
3Claude Opus 4.573.25
4Kimi K2.571.84
5Gemini 3 Flash (Preview)71.62
6DeepSeek V3.270.71
7DeepSeek R170.58
8GLM-4.769.64
9Gemini 2.5 Pro69.6
10Seed 1.869.03
11DeepSeek V367.57
12Qwen 3 32B66.71
13Qwen 2.5 72B Instruct65.78
14GPT-4.165.46
15Qwen 3 14B64.81

Interactive version: theaggregate.ai/benchmark?slug=sage-service-agent-property-service · How It Works · Data refreshed daily, snapshot 2026-10-07.