HarmThoughts (Few-shot): leaderboard
Metric: Macro-F1 (x100) of 16-way sentence-level behavior classification (harm-propagation, safety-preserving and neutral behaviors) on jailbroken reasoning traces of HarmThoughts, the model prompted with the taxonomy and the full trace, few-shot prompting with human-annotated examples; scored over the parseable predictions; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 56.21 |
| 2 | Gemini 2.0 Flash | 52.52 |
| 3 | GPT-4o | 40.85 |
| 4 | Llama 3.3 70B | 40.41 |
Interactive version: theaggregate.ai/benchmark?slug=harmthoughts-few-shot · How It Works · Data refreshed daily, snapshot 2026-10-07.