HuluMed-7B: benchmark results
Provider: Other.
Unified ELO 1528 ± 29, rank #624 of 1537 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CareQA-Vision - Open-Ended - Medicine | 48.61 | Judge score (%; 108 open-ended image-based questions from MI | 64.7 |
| CareQA-Vision - Open-Ended | 43.38 | Judge score (%; 136 open-ended image-based questions from MI | 61.1 |
| CareQA-Vision - MCQ - Medicine | 65.85 | Accuracy (%; 123 image-based multiple-choice questions from | 52.9 |
| MedPIC-Bench - Rule Deactivation | 38.2 | Accuracy (%) | 51.9 |
| CareQA-Vision - Open-Ended - Nursing | 23.21 | Judge score (%; 28 open-ended image-based questions from EIR | 50 |
| CareQA-Vision - MCQ | 59.39 | Accuracy (%; 165 image-based multiple-choice questions from | 47.2 |
| MedPIC-Bench - Counterfactual Pair | 14.6 | Both-correct pair rate (%) | 38.9 |
| BRIDGE Medical Leaderboard - Few-Shot | 40.6 | Average Performance (%) | 37 |
| BRIDGE Medical Leaderboard | 32.97 | Average Performance (%) | 33.3 |
| BRIDGE Medical Leaderboard - Zero-Shot | 30.71 | Average Performance (%) | 33.3 |
| BRIDGE Medical Leaderboard - CoT | 27.6 | Average Performance (%) | 31.5 |
| CardioLens | 44.81 | F1 (Random) (self-reported) | 30.4 |
Interactive version: theaggregate.ai/model?slug=hulumed-7b · How It Works · Data refreshed daily, snapshot 2026-09-25.