Medical STT Benchmark — leaderboard

Speech-to-text benchmark for long-form medical dialogue, ranking cloud and local transcription systems with Medical Word Error Rate on the PriMock57 dataset.

Metric: WER (lower is better). Source: github.com. Status: saturation imminent. 42 models tracked.

Top models

#ModelScore
1gemma-4-E4B-it0.16
2gemma-4-E2B-it0.19

Interactive version: theaggregate.ai/benchmark?slug=medical-stt-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.