MedMosaic - Open-Ended Speech: leaderboard
Metric: Judge score (x 100; GPT-5.1 rates correctness, relevance, completeness and clarity against the reference answer on 0-1 scales, unweighted mean; 5,706 open-ended questions on clinical speech). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 79.2 |
| 2 | Gemini 2.5 Flash | 74.8 |
| 3 | Gemma 3n E4B (IT) | 64.8 |
| 4 | GPT-4o Audio | 60.8 |
| 5 | Qwen2.5-Omni-7B | 57 |
| 6 | Phi-4 Multimodal Instruct | 54.3 |
Interactive version: theaggregate.ai/benchmark?slug=medmosaic-open-ended-speech · How It Works · Data refreshed daily, snapshot 2026-09-26.