HINTBench - Coarse Risk-Step Localization (Strict): leaderboard
Metric: Strict F1 (%) for coarse risk-step localization: a predicted risky step matches a gold step within plus or minus 3 steps and the predicted primary constraint category must match the gold one, over the 536 synthetic HINTBench agent trajectories (400 risky, 136 safe, 24 steps on average, unanimously accepted by three human verifiers) audited zero-shot by a general LLM for intrinsic, non-attack risk under a five-constraint taxonomy (goal, capability, factual, procedural, state constraints), each model at its stated reasoning setting; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (High) | 45.96 |
| 2 | Claude Sonnet 4.6 (Non-reasoning) | 45.66 |
| 3 | GPT-5.5 (Low) | 44.76 |
| 4 | Claude Opus 4.6 (Non-reasoning) | 44.57 |
| 5 | GPT-5.5 (Medium) | 44.4 |
| 6 | Kimi K2.5 | 43.63 |
| 7 | GLM-5 | 41.19 |
| 8 | ERNIE 5.0 | 40.67 |
| 9 | DeepSeek R1 | 40.25 |
| 10 | MiniMax-M2.5 | 39.88 |
| 11 | Qwen 3 235B A22B | 36.28 |
| 12 | DeepSeek V3.2 | 31.93 |
| 13 | Qwen 3 32B | 23.2 |
| 14 | Qwen 3 14B | 22.87 |
| 15 | Qwen 3 8B | 14.34 |
Interactive version: theaggregate.ai/benchmark?slug=hintbench-coarse-risk-step-localization-strict · How It Works · Data refreshed daily, snapshot 2026-10-07.