KC-Bench - Region: leaderboard
Metric: Success rate (%; world-knowledge conflict: 71 region tasks in which the user asserts a false fact and the agent must reject the premise before using read-only factual tools; one trial per task, greedy decoding, at most 30 interaction steps, LLM user simulator, stateful tools, success only if the conflict is detected and no protected state is exposed or modified; human-verified label (automatic assertion verdicts adjudicated by annotators)). Source: arxiv.org. Saturation forecast: Around September 2028. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | MiniMax-M3 | 67.61 |
| 2 | GLM-5.2 | 53.52 |
| 3 | Qwen 3.5 35B A3B | 50.7 |
| 4 | Claude Haiku 4.5 (20251001) | 45.07 |
| 5 | GPT-5.1 | 40.85 |
| 6 | GLM 4.5 Air | 38.04 |
| 7 | DeepSeek V4 Flash | 18.31 |
| 8 | GPT-OSS-120B | 8.45 |
| 9 | Gemini 3 Flash (Preview) | 7.04 |
Interactive version: theaggregate.ai/benchmark?slug=kc-bench-region · How It Works · Data refreshed daily, snapshot 2026-09-26.