MedRealMM (Physician Subset): leaderboard

Metric: Rubric score (%; case-level mean of the physician-refined case rubric, positive and negative criteria weighted and clipped to 0-1, graded per criterion by Gemini-3-Pro-Preview) on the 200-case subset of real Chinese online consultations with patient images used for the deliberative physician comparison. Source: arxiv.org. Saturation forecast: Around 2028. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.750.66
2Gemini 3.5 Flash44.93
3Kimi K2.641.78
4Qwen 3.6 27B41.2
5GPT-5.538.53

Interactive version: theaggregate.ai/benchmark?slug=medrealmm-physician-subset · How It Works · Data refreshed daily, snapshot 2026-09-29.