Hulu-Med-14B: benchmark results

Provider: Other.

Unified ELO 1540 ± 27, rank #901 of 2131 rated models, from 26 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Med-R2 - Anatomy Level68.1Accuracy (%) on the Med-R2 Anatomy Level task (locating and 100
Med-R2 - Lesion Level68.6Accuracy (%) on the Med-R2 Lesion Level task (detecting and 100
Med-R261.8Overall average accuracy (%) over the Image Quality, Anatomy88.5
MMRareBench - Diagnosis52Diagnosis track (T1) score: open-ended primary diagnosis fro63.6
MMRareBench - Diagnosis F158.2Diagnosis track (T1) token-level F1 (%, SQuAD-style normaliz63.6
Med-R2 - Joint QA63.1Accuracy (%) on the Med-R2 Joint QA task (answering all reas61.5
BRIDGE Medical Leaderboard - Few-Shot45.87Average Performance (%)60.4
MMRareBench - Treatment Planning14.2Treatment Planning track (T2) score: a stage-wise treatment 52.3
BRIDGE Medical Leaderboard - Zero-Shot35.45Average Performance (%)50.5
BRIDGE Medical Leaderboard - CoT31.61Average Performance (%)45
MMRareBench - Examination Suggestion38.8Examination Suggestion track (T4) score: prioritized next-st40.9
MMRareBench - Cross-Image Evidence Alignment6.1Cross-Image Evidence Alignment track (T3) score: per-image f31.8

Interactive version: theaggregate.ai/model?slug=hulu-med-14b · How It Works · Data refreshed daily, snapshot 2026-10-09.