Hulu-Med-32B: benchmark results

Provider: Other.

Unified ELO 1576 ± 26, rank #516 of 1629 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CareQA-Vision - Open-Ended - Medicine53.7Judge score (%; 108 open-ended image-based questions from MI94.1
CheXpercept91.4Stage 1 (End-to-End) (self-reported)92.3
CareQA-Vision - Open-Ended48.53Judge score (%; 136 open-ended image-based questions from MI80.6
BRIDGE Medical Leaderboard - Few-Shot48.3Average Performance (%)72.5
CareQA-Vision - MCQ - Medicine72.36Accuracy (%; 123 image-based multiple-choice questions from 70.6
CareQA-Vision - Open-Ended - Nursing28.57Judge score (%; 28 open-ended image-based questions from EIR70.6
BRIDGE Medical Leaderboard40.28Average Performance (%)64
BRIDGE Medical Leaderboard - Zero-Shot38.29Average Performance (%)64
CareQA-Vision - MCQ63.64Accuracy (%; 165 image-based multiple-choice questions from 61.1
BRIDGE Medical Leaderboard - CoT34.25Average Performance (%)55.9
JMed48k (With Images)36.8Accuracy (%) on the 2,579 scored JMed48k-Eval items that inc31.6
JMed48k (Text-Only) - Dentist56.4Accuracy (%) on the text-only scored items of the Japanese D30

Interactive version: theaggregate.ai/model?slug=hulu-med-32b · How It Works · Data refreshed daily, snapshot 2026-10-07.