MedRepBench - Explanation Acceptability: leaderboard

Metric: Acceptability rate (%; share of patient-facing explanations a DeepSeek-R1 judge scores 1 or 2 on a 0-2 scale for factuality, reasoning and clarity; VLM on the report image only). Source: arxiv.org. Saturation forecast: Estimated already saturated. 9 models tracked.

Top models

#ModelScore
1Llama 4 Maverick60.78
2InternVL3-38B51.15
3InternVL3-8B48.24
4InternVL3-14B47
5Llama 4 Scout46.84
6Qwen 2.5 VL 7B26.21

Interactive version: theaggregate.ai/benchmark?slug=medrepbench-explanation-acceptability · How It Works · Data refreshed daily, snapshot 2026-09-26.