KC-Bench - Retail: leaderboard

Metric: Success rate (%; input inconsistency: 93 retail customer-service tasks in which the user credentials match a database record on email but not name, and the agent must clarify before disclosing or changing account state; one trial per task, greedy decoding, at most 30 interaction steps, LLM user simulator, stateful tools, success only if the conflict is detected and no protected state is exposed or modified; human-verified label (automatic assertion verdicts adjudicated by annotators)). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Claude Haiku 4.5 (20251001)73.12
2DeepSeek V4 Flash65.22
3GLM-5.261.96
4GPT-5.150.54
5Gemini 3 Flash (Preview)50.54
6GLM 4.5 Air46.23
7MiniMax-M338.04
8Qwen 3.5 35B A3B19.35
9GPT-OSS-120B11.82

Interactive version: theaggregate.ai/benchmark?slug=kc-bench-retail · How It Works · Data refreshed daily, snapshot 2026-09-26.