PlaceboBench — leaderboard

Medical-domain hallucination benchmark with labeled model answers to pharmaceutical questions grounded in EMA product information.

Metric: Non-Hallucination Rate (self-reported). Source: benchmarklist.com. 7 models tracked.

Top models

#ModelScore
1GPT-5.263.24
2Claude Sonnet 4.562.32
3Gemini 3 Flash (Preview)44.93
4GPT-5 Mini39.13
5Claude Opus 4.636.23

Interactive version: theaggregate.ai/benchmark?slug=placebobench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.