MedMosaic: leaderboard

Metric: Weighted accuracy (%; category scores weighted by question count over 46,821 scored questions, the 60 multi-turn chains counted as 180 turns; open-ended categories scored by a GPT-5.1 judge). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro68.1
2Gemini 2.5 Flash60.5
3Qwen2.5-Omni-7B42.8
4Gemma 3n E4B (IT)42.1
5Phi-4 Multimodal Instruct37.3
6GPT-4o Audio35.7

Interactive version: theaggregate.ai/benchmark?slug=medmosaic · How It Works · Data refreshed daily, snapshot 2026-09-26.