Omi SOAP Note Safety Benchmark — leaderboard

Safety-first clinical SOAP note generation benchmark measuring groundedness, hallucinations, coverage, and note quality across 300 doctor-patient dialogues.

Metric: Composite (self-reported). Source: benchmarklist.com. 6 models tracked.

Top models

#ModelScore
1GPT-5.24.72
2Kimi K2 (Thinking)4.55
3Claude Opus 4.54.54
4GPT-54.29

Interactive version: theaggregate.ai/benchmark?slug=omi-soap-note-safety-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.