CarryOnBench - Utility (Oracle Intent): leaderboard

Metric: Ben-Util (%): share of the benign information-need checklist (about 18 atomic yes/no items per query, built from five models' compliant answers) that the responses fulfil, each item judged separately by Gemini-2.5-Flash, over a single turn in which the query states the full benign intent, on CarryOnBench (398 seemingly harmful SORRY-Bench queries with a benign underlying intent; user turns simulated by Gemini-3-Flash); higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 14 models tracked.

Top models

#ModelScore
1GPT-5 Mini72.1
2Qwen 3 Next 80B A3B Instruct70.7
3DeepSeek V3.167.6
4GPT-5.463.2
5Qwen 3 32B61.2
6Qwen 3 235B A22B56.8
7Olmo 3.1 32B Instruct56.8
8Gemini 3.1 Pro (Preview)52.1
9Claude Sonnet 448.8
10OLMo 3 7B Instruct48.8
11Claude Haiku 4.545.9
12Mixtral 8x7B Instruct (v0.1)40
13Llama 4 Maverick35.7
14GPT-OSS-120B25.2

Interactive version: theaggregate.ai/benchmark?slug=carryonbench-utility-oracle-intent · How It Works · Data refreshed daily, snapshot 2026-10-07.