HINTBench - Fine Risk-Step Localization (Strict): leaderboard
Metric: Strict F1 (%) for fine risk-step localization: a predicted risky step matches a gold step within plus or minus 3 steps and the predicted primary subtype must match the gold one, over the 536 synthetic HINTBench agent trajectories (400 risky, 136 safe, 24 steps on average, unanimously accepted by three human verifiers) audited zero-shot by a general LLM for intrinsic, non-attack risk under a five-constraint taxonomy (goal, capability, factual, procedural, state constraints), each model at its stated reasoning setting; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Non-reasoning) | 37.93 |
| 2 | Claude Sonnet 4.6 (Non-reasoning) | 36.84 |
| 3 | GLM-5 | 36.01 |
| 4 | GPT-5.5 (Low) | 34.8 |
| 5 | GPT-5.5 (High) | 34.17 |
| 6 | Kimi K2.5 | 33.71 |
| 7 | GPT-5.5 (Medium) | 33.58 |
| 8 | DeepSeek R1 | 32.76 |
| 9 | ERNIE 5.0 | 32.64 |
| 10 | Qwen 3 235B A22B | 31.85 |
| 11 | MiniMax-M2.5 | 28.34 |
| 12 | DeepSeek V3.2 | 27.27 |
| 13 | Qwen 3 14B | 26.77 |
| 14 | Qwen 3 32B | 25.39 |
| 15 | Qwen 3 8B | 10.67 |
Interactive version: theaggregate.ai/benchmark?slug=hintbench-fine-risk-step-localization-strict · How It Works · Data refreshed daily, snapshot 2026-10-07.