CarryOnBench - Utility (Benign Clarifications): leaderboard

Metric: Ben-Util (%): share of the benign information-need checklist (about 18 atomic yes/no items per query, built from five models' compliant answers) that the responses fulfil, each item judged separately by Gemini-2.5-Flash, over multi-turn conversations of 4 to 12 turns in which at least one user turn clarifies the benign intent, conversation-level, on CarryOnBench (398 seemingly harmful SORRY-Bench queries with a benign underlying intent; user turns simulated by Gemini-3-Flash); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 14 models tracked.

Top models

#ModelScore
1Qwen 3 Next 80B A3B Instruct75.2
2GPT-5 Mini74.8
3Qwen 3 32B73.8
4Qwen 3 235B A22B71.2
5DeepSeek V3.166.3
6GPT-5.465.4
7Olmo 3.1 32B Instruct64.4
8OLMo 3 7B Instruct61.3
9Gemini 3.1 Pro (Preview)57.4
10GPT-OSS-120B52.5
11Claude Sonnet 451.3
12Claude Haiku 4.549.9
13Mixtral 8x7B Instruct (v0.1)45.1
14Llama 4 Maverick37

Interactive version: theaggregate.ai/benchmark?slug=carryonbench-utility-benign-clarifications · How It Works · Data refreshed daily, snapshot 2026-10-07.