WHBench - Harm Rate: leaderboard

Metric: Harm rate (%): share of responses failing either the unsafe-commission or the unsafe-omission safety criterion on the 47 expert-written WHBench scenarios in ten women's health topics (fertility, hormonal health, pregnancy, PCOS, contraception and others), three runs per question at temperature 0 with a 4,096-token cap, each response graded by a Claude Sonnet 4.6 judge against 4-6 clinician reference answers and a 23-criterion rubric with asymmetric safety penalties; lower is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Claude Opus 4.612.8
2Claude Sonnet 4.627
3Mistral Large 330.5
4Gemini 3 Flash (Preview)32.6
5Grok 333.3
6Grok 437.1
7O338.6
8DeepSeek V3.244
9GPT-5.447.5
10DeepSeek R147.5
11Grok 3 Mini53.9
12Claude Opus 456
13GPT-4.161
14Claude Sonnet 468.1
15Gemini 2.5 Flash73.8

Interactive version: theaggregate.ai/benchmark?slug=whbench-harm-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.