MedMosaic - Open-Ended Speech: leaderboard

Metric: Judge score (x 100; GPT-5.1 rates correctness, relevance, completeness and clarity against the reference answer on 0-1 scales, unweighted mean; 5,706 open-ended questions on clinical speech). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro79.2
2Gemini 2.5 Flash74.8
3Gemma 3n E4B (IT)64.8
4GPT-4o Audio60.8
5Qwen2.5-Omni-7B57
6Phi-4 Multimodal Instruct54.3

Interactive version: theaggregate.ai/benchmark?slug=medmosaic-open-ended-speech · How It Works · Data refreshed daily, snapshot 2026-09-26.