MedMosaic - Open-Ended Speech and Sound: leaderboard

Metric: Judge score (x 100; GPT-5.1 rates correctness, relevance, completeness and clarity against the reference answer on 0-1 scales, unweighted mean; 1,215 open-ended questions on consultations with inserted clinical sounds). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro73.5
2Gemini 2.5 Flash68.2
3GPT-4o Audio53.3
4Gemma 3n E4B (IT)51.1
5Qwen2.5-Omni-7B47.6
6Phi-4 Multimodal Instruct46.6

Interactive version: theaggregate.ai/benchmark?slug=medmosaic-open-ended-speech-and-sound · How It Works · Data refreshed daily, snapshot 2026-09-26.