HELM AIR-Bench 2024 - #14: Hate Speech (Inciting/Promoting/Expressing hatred): leaderboard

Metric: Refusal Rate (%). Source: crfm.stanford.edu. 87 models tracked.

Top models

#ModelScore
1Claude Haiku 4.5 (20251001)100
2Claude Sonnet 4 (20250514)99.86
3Claude Sonnet 4.599.71
4Claude Opus 4 (20250514)99.71
5Claude 3.5 Sonnet (20241022)99.57
6Qwen 3 Next 80B A3B (Thinking)98.99
7Claude 3.7 Sonnet (20250219)98.99
8Claude 3.5 Sonnet (20240620)98.55
9Claude 3 Opus (20240229)98.12
10GPT-OSS-20B97.97
11Claude 3 Haiku (20240307)97.97
12GPT-597.83
13GPT-5 Nano97.68
14GPT-OSS-120B97.68
15Claude 3 Sonnet (20240229)97.68

Interactive version: theaggregate.ai/benchmark?slug=helm-air-bench-2024-14-hate-speech-inciting-promoting-expressing-hatred · How It Works · Data refreshed daily, snapshot 2026-09-19.