HarmThoughts (Zero-shot): leaderboard

Metric: Macro-F1 (x100) of 16-way sentence-level behavior classification (harm-propagation, safety-preserving and neutral behaviors) on jailbroken reasoning traces of HarmThoughts, the model prompted with the taxonomy and the full trace, zero-shot prompting; scored over the parseable predictions; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro50.97
2Gemini 2.0 Flash40.88
3GPT-4o40.72
4Llama 3.3 70B33.05

Interactive version: theaggregate.ai/benchmark?slug=harmthoughts-zero-shot · How It Works · Data refreshed daily, snapshot 2026-10-07.