MedMosaic - Open-Ended Speech and Sound: leaderboard
Metric: Judge score (x 100; GPT-5.1 rates correctness, relevance, completeness and clarity against the reference answer on 0-1 scales, unweighted mean; 1,215 open-ended questions on consultations with inserted clinical sounds). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 73.5 |
| 2 | Gemini 2.5 Flash | 68.2 |
| 3 | GPT-4o Audio | 53.3 |
| 4 | Gemma 3n E4B (IT) | 51.1 |
| 5 | Qwen2.5-Omni-7B | 47.6 |
| 6 | Phi-4 Multimodal Instruct | 46.6 |
Interactive version: theaggregate.ai/benchmark?slug=medmosaic-open-ended-speech-and-sound · How It Works · Data refreshed daily, snapshot 2026-09-26.