S2SServiceBench (Signal Comprehension): leaderboard

Metric: Level 1 overall score (0-1 printed, shown times 100): unweighted mean over the ten service products (161 expert-selected cases in all) of short structured fields (booleans, numbers within 5%, and categories or regions judged by an LLM) extracted from the product, scored per field; direct prompting with the product artifacts, one response per task; GPT-5.2 is also a subject; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 6 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.236.13#105
2Gemini 3 Pro36.04#77
3Claude Opus 4.527.33#79
4Llama 4 Maverick Instruct26.1#439
5Qwen 3 VL 32B Instruct25.91#276
6Grok 419.96#169

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=s2sservicebench-signal-comprehension · How It Works · Data refreshed daily, snapshot 2026-10-11.