KC-Bench - Retail: leaderboard
Metric: Success rate (%; input inconsistency: 93 retail customer-service tasks in which the user credentials match a database record on email but not name, and the agent must clarify before disclosing or changing account state; one trial per task, greedy decoding, at most 30 interaction steps, LLM user simulator, stateful tools, success only if the conflict is detected and no protected state is exposed or modified; human-verified label (automatic assertion verdicts adjudicated by annotators)). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 (20251001) | 73.12 |
| 2 | DeepSeek V4 Flash | 65.22 |
| 3 | GLM-5.2 | 61.96 |
| 4 | GPT-5.1 | 50.54 |
| 5 | Gemini 3 Flash (Preview) | 50.54 |
| 6 | GLM 4.5 Air | 46.23 |
| 7 | MiniMax-M3 | 38.04 |
| 8 | Qwen 3.5 35B A3B | 19.35 |
| 9 | GPT-OSS-120B | 11.82 |
Interactive version: theaggregate.ai/benchmark?slug=kc-bench-retail · How It Works · Data refreshed daily, snapshot 2026-09-26.