PlaceboBench: leaderboard

Medical-domain hallucination benchmark with labeled model answers to pharmaceutical questions grounded in EMA product information.

Metric: Non-Hallucination Rate (self-reported). Source: benchmarklist.com. 7 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)73.91
2GPT-5.263.24
3Claude Sonnet 4.562.32
4Gemini 3 Flash (Preview)44.93
5GPT-5 Mini39.13
6Claude Opus 4.636.23

Interactive version: theaggregate.ai/benchmark?slug=placebobench · How It Works · Data refreshed daily, snapshot 2026-09-05.