HINTBench - Coarse Risk-Step Localization (Strict): leaderboard

Metric: Strict F1 (%) for coarse risk-step localization: a predicted risky step matches a gold step within plus or minus 3 steps and the predicted primary constraint category must match the gold one, over the 536 synthetic HINTBench agent trajectories (400 risky, 136 safe, 24 steps on average, unanimously accepted by three human verifiers) audited zero-shot by a general LLM for intrinsic, non-attack risk under a five-constraint taxonomy (goal, capability, factual, procedural, state constraints), each model at its stated reasoning setting; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 21 models tracked.

Top models

#ModelScore
1GPT-5.5 (High)45.96
2Claude Sonnet 4.6 (Non-reasoning)45.66
3GPT-5.5 (Low)44.76
4Claude Opus 4.6 (Non-reasoning)44.57
5GPT-5.5 (Medium)44.4
6Kimi K2.543.63
7GLM-541.19
8ERNIE 5.040.67
9DeepSeek R140.25
10MiniMax-M2.539.88
11Qwen 3 235B A22B36.28
12DeepSeek V3.231.93
13Qwen 3 32B23.2
14Qwen 3 14B22.87
15Qwen 3 8B14.34

Interactive version: theaggregate.ai/benchmark?slug=hintbench-coarse-risk-step-localization-strict · How It Works · Data refreshed daily, snapshot 2026-10-07.