Hulu-Med-14B: benchmark results
Provider: Other.
Unified ELO 1540 ± 27, rank #901 of 2131 rated models, from 26 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Med-R2 - Anatomy Level | 68.1 | Accuracy (%) on the Med-R2 Anatomy Level task (locating and | 100 |
| Med-R2 - Lesion Level | 68.6 | Accuracy (%) on the Med-R2 Lesion Level task (detecting and | 100 |
| Med-R2 | 61.8 | Overall average accuracy (%) over the Image Quality, Anatomy | 88.5 |
| MMRareBench - Diagnosis | 52 | Diagnosis track (T1) score: open-ended primary diagnosis fro | 63.6 |
| MMRareBench - Diagnosis F1 | 58.2 | Diagnosis track (T1) token-level F1 (%, SQuAD-style normaliz | 63.6 |
| Med-R2 - Joint QA | 63.1 | Accuracy (%) on the Med-R2 Joint QA task (answering all reas | 61.5 |
| BRIDGE Medical Leaderboard - Few-Shot | 45.87 | Average Performance (%) | 60.4 |
| MMRareBench - Treatment Planning | 14.2 | Treatment Planning track (T2) score: a stage-wise treatment | 52.3 |
| BRIDGE Medical Leaderboard - Zero-Shot | 35.45 | Average Performance (%) | 50.5 |
| BRIDGE Medical Leaderboard - CoT | 31.61 | Average Performance (%) | 45 |
| MMRareBench - Examination Suggestion | 38.8 | Examination Suggestion track (T4) score: prioritized next-st | 40.9 |
| MMRareBench - Cross-Image Evidence Alignment | 6.1 | Cross-Image Evidence Alignment track (T3) score: per-image f | 31.8 |
Interactive version: theaggregate.ai/model?slug=hulu-med-14b · How It Works · Data refreshed daily, snapshot 2026-10-09.