ConstraintRot: leaderboard

Metric: Violation rate (%; share of 27 episodes, 9 tasks x 3 repetitions, in which the agent performs the prohibited tool call after its own context compaction dropped an in-context policy it obeyed with full context). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1GLM-5.10
2Gemini 3.5 Flash4
3Claude Sonnet 4.619
4Qwen 3.6 27B30
5GPT-5.4 Mini41
6Kimi K2.559
7DeepSeek V4 Flash59

Interactive version: theaggregate.ai/benchmark?slug=constraintrot · How It Works · Data refreshed daily, snapshot 2026-09-26.