S2SServiceBench (Decision-Making Handoff): leaderboard

Metric: Level 2 overall score (0-1 printed, shown times 100): unweighted mean over the ten service products (161 expert-selected cases in all) of a schema-constrained handoff report with triggers, timing, constraints and uncertainty-conditioned actions, graded 0-5 on six rubric dimensions by a GPT-5.2 judge and printed on a 0-1 scale; direct prompting with the product artifacts, one response per task; GPT-5.2 is also a subject; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 6 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.263.79#105
2Claude Opus 4.550.54#79
3Gemini 3 Pro37.92#77
4Grok 436.71#169
5Qwen 3 VL 32B Instruct35.22#276
6Llama 4 Maverick Instruct20.14#439

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=s2sservicebench-decision-making-handoff · How It Works · Data refreshed daily, snapshot 2026-10-11.