TrustMH-Bench - Jailbreak Refusal Rate: leaderboard
Metric: Overall refusal rate (%) on JailbreakMH: 560 harmful mental-health requests (70 under each of eight jailbreak methods such as prefix injection, refusal suppression and multi-task interference), a refusal judged by the LibrAI longformer-harmful-ro classifier; temperature 0; a safety propensity, higher is better. Source: arxiv.org. 12 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.1 | 98.7 | #131 |
| 2 | Claude Sonnet 4.5 | 97 | #138 |
| 3 | Gemini 2.5 Flash | 90 | #237 |
| 4 | Qwen 3 235B A22B 2507 Instruct | 88.2 | #291 |
| 5 | GPT-4o Mini | 82.5 | #588 |
| 6 | DeepSeek V3.2 | 75.2 | #198 |
Interactive version: theaggregate.ai/benchmark?slug=trustmh-bench-jailbreak-refusal-rate · How It Works · Data refreshed daily, snapshot 2026-10-11.