Hulu-Med-32B: benchmark results
Provider: Other.
Unified ELO 1576 ± 26, rank #516 of 1629 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CareQA-Vision - Open-Ended - Medicine | 53.7 | Judge score (%; 108 open-ended image-based questions from MI | 94.1 |
| CheXpercept | 91.4 | Stage 1 (End-to-End) (self-reported) | 92.3 |
| CareQA-Vision - Open-Ended | 48.53 | Judge score (%; 136 open-ended image-based questions from MI | 80.6 |
| BRIDGE Medical Leaderboard - Few-Shot | 48.3 | Average Performance (%) | 72.5 |
| CareQA-Vision - MCQ - Medicine | 72.36 | Accuracy (%; 123 image-based multiple-choice questions from | 70.6 |
| CareQA-Vision - Open-Ended - Nursing | 28.57 | Judge score (%; 28 open-ended image-based questions from EIR | 70.6 |
| BRIDGE Medical Leaderboard | 40.28 | Average Performance (%) | 64 |
| BRIDGE Medical Leaderboard - Zero-Shot | 38.29 | Average Performance (%) | 64 |
| CareQA-Vision - MCQ | 63.64 | Accuracy (%; 165 image-based multiple-choice questions from | 61.1 |
| BRIDGE Medical Leaderboard - CoT | 34.25 | Average Performance (%) | 55.9 |
| JMed48k (With Images) | 36.8 | Accuracy (%) on the 2,579 scored JMed48k-Eval items that inc | 31.6 |
| JMed48k (Text-Only) - Dentist | 56.4 | Accuracy (%) on the text-only scored items of the Japanese D | 30 |
Interactive version: theaggregate.ai/model?slug=hulu-med-32b · How It Works · Data refreshed daily, snapshot 2026-10-07.