KC-Bench - Region: leaderboard

Metric: Success rate (%; world-knowledge conflict: 71 region tasks in which the user asserts a false fact and the agent must reject the premise before using read-only factual tools; one trial per task, greedy decoding, at most 30 interaction steps, LLM user simulator, stateful tools, success only if the conflict is detected and no protected state is exposed or modified; human-verified label (automatic assertion verdicts adjudicated by annotators)). Source: arxiv.org. Saturation forecast: Around September 2028. 9 models tracked.

Top models

#ModelScore
1MiniMax-M367.61
2GLM-5.253.52
3Qwen 3.5 35B A3B50.7
4Claude Haiku 4.5 (20251001)45.07
5GPT-5.140.85
6GLM 4.5 Air38.04
7DeepSeek V4 Flash18.31
8GPT-OSS-120B8.45
9Gemini 3 Flash (Preview)7.04

Interactive version: theaggregate.ai/benchmark?slug=kc-bench-region · How It Works · Data refreshed daily, snapshot 2026-09-26.