Medical STT Benchmark — leaderboard
Speech-to-text benchmark for long-form medical dialogue, ranking cloud and local transcription systems with Medical Word Error Rate on the PriMock57 dataset.
Metric: WER (lower is better). Source: github.com. Status: saturation imminent. 42 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gemma-4-E4B-it | 0.16 |
| 2 | gemma-4-E2B-it | 0.19 |
Interactive version: theaggregate.ai/benchmark?slug=medical-stt-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.