S2SServiceBench (Decision Analysis and Planning): leaderboard

Metric: Level 3 overall score (0-1 printed, shown times 100): unweighted mean over the ten service products (161 expert-selected cases in all) of a schema-constrained strategic assessment and planning report, graded 0-5 on six rubric dimensions by a GPT-5.2 judge and printed on a 0-1 scale; direct prompting with the product artifacts, one response per task; GPT-5.2 is also a subject; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 6 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.250.88#105
2Claude Opus 4.541.81#79
3Gemini 3 Pro32.68#77
4Qwen 3 VL 32B Instruct31.42#276
5Grok 431.34#169
6Llama 4 Maverick Instruct20.16#439

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=s2sservicebench-decision-analysis-and-planning · How It Works · Data refreshed daily, snapshot 2026-10-11.