Omi SOAP Note Safety Benchmark: leaderboard

Safety-first clinical SOAP note generation benchmark measuring groundedness, hallucinations, coverage, and note quality across 300 doctor-patient dialogues.

Metric: Composite (self-reported). Source: benchmarklist.com. 6 models tracked.

Top models

#ModelScore
1GPT-5.24.72
2Gemini 3 Pro (Preview)4.7
3Kimi K2 (Thinking)4.55
4Claude Opus 4.54.54
5GPT-54.29

Interactive version: theaggregate.ai/benchmark?slug=omi-soap-note-safety-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-05.