HINTBench - Fine Risk-Step Localization (Strict): leaderboard

Metric: Strict F1 (%) for fine risk-step localization: a predicted risky step matches a gold step within plus or minus 3 steps and the predicted primary subtype must match the gold one, over the 536 synthetic HINTBench agent trajectories (400 risky, 136 safe, 24 steps on average, unanimously accepted by three human verifiers) audited zero-shot by a general LLM for intrinsic, non-attack risk under a five-constraint taxonomy (goal, capability, factual, procedural, state constraints), each model at its stated reasoning setting; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Non-reasoning)37.93
2Claude Sonnet 4.6 (Non-reasoning)36.84
3GLM-536.01
4GPT-5.5 (Low)34.8
5GPT-5.5 (High)34.17
6Kimi K2.533.71
7GPT-5.5 (Medium)33.58
8DeepSeek R132.76
9ERNIE 5.032.64
10Qwen 3 235B A22B31.85
11MiniMax-M2.528.34
12DeepSeek V3.227.27
13Qwen 3 14B26.77
14Qwen 3 32B25.39
15Qwen 3 8B10.67

Interactive version: theaggregate.ai/benchmark?slug=hintbench-fine-risk-step-localization-strict · How It Works · Data refreshed daily, snapshot 2026-10-07.