HarmThoughts (Few-shot): leaderboard

Metric: Macro-F1 (x100) of 16-way sentence-level behavior classification (harm-propagation, safety-preserving and neutral behaviors) on jailbroken reasoning traces of HarmThoughts, the model prompted with the taxonomy and the full trace, few-shot prompting with human-annotated examples; scored over the parseable predictions; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro56.21
2Gemini 2.0 Flash52.52
3GPT-4o40.85
4Llama 3.3 70B40.41

Interactive version: theaggregate.ai/benchmark?slug=harmthoughts-few-shot · How It Works · Data refreshed daily, snapshot 2026-10-07.