WHBench: leaderboard

Metric: Mean normalized rubric score (%), the raw score from -58 to +92 mapped to 0-100 on the 47 expert-written WHBench scenarios in ten women's health topics (fertility, hormonal health, pregnancy, PCOS, contraception and others), three runs per question at temperature 0 with a 4,096-token cap, each response graded by a Claude Sonnet 4.6 judge against 4-6 clinician reference answers and a 23-criterion rubric with asymmetric safety penalties; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Claude Opus 4.672.1
2Claude Sonnet 4.667.1
3GPT-5.466.8
4Gemini 3 Flash (Preview)64.7
5O363.6
6DeepSeek V3.261.3
7Grok 360.7
8Mistral Large 360.2
9Grok 457.9
10DeepSeek R152.9
11GPT-4.151.8
12Grok 3 Mini50
13Gemini 2.5 Flash49.5
14Claude Opus 449.1
15Claude Sonnet 448.1

Interactive version: theaggregate.ai/benchmark?slug=whbench · How It Works · Data refreshed daily, snapshot 2026-10-07.