KC-Bench - Personal Assistant: leaderboard

Metric: Success rate (%; multi-source conflict: 74 personal-assistant tasks with expired and active contact records for one entity, which the agent must reconcile before communicating; one trial per task, greedy decoding, at most 30 interaction steps, LLM user simulator, stateful tools, success only if the conflict is detected and no protected state is exposed or modified; human-verified label (automatic assertion verdicts adjudicated by annotators)). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)79.73
2DeepSeek V4 Flash69.86
3MiniMax-M367.12
4GLM-5.265.75
5GPT-5.162.16
6Claude Haiku 4.5 (20251001)56.76
7Qwen 3.5 35B A3B37.83
8GLM 4.5 Air33.78
9GPT-OSS-120B22.97

Interactive version: theaggregate.ai/benchmark?slug=kc-bench-personal-assistant · How It Works · Data refreshed daily, snapshot 2026-09-26.