Llama 3 8B UltraMedical: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1414 ± 29, rank #1164 of 1639 rated models, from 16 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - MMLU66.17Accuracy (%) (5-shot)86.1
CureMed-Bench - Language Consistency47.03Language consistency (%; share of final answers written in t64.6
Open LLM Leaderboard v1 - GSM8K44.81Accuracy (%) (5-shot)58.8
Open LLM Leaderboard v1 - TruthfulQA MC252.32MC2 (%) (0-shot)55.4
Open LLM Leaderboard v1 - ARC Challenge61.52Normalized accuracy (%) (25-shot)53.2
Open LLM Leaderboard v1 - HellaSwag82.37Normalized accuracy (%) (10-shot)50.8
CureMed-Bench35.29Logical accuracy (%; share of final answers a GPT-4o judge m45.8
Open LLM Leaderboard v1 - WinoGrande75.22Accuracy (%) (5-shot)40.9
MedRoundsQA - Multi-Turn Diagnosis22.21Accuracy (%)28.6
MedRoundsQA - Single-Turn Diagnosis50Accuracy (%)28.6
EHRBench - Treatment43.09Accuracy (%) on the treatment-decision questions, EHR-ground3.3
EHRBench - Diagnosis19.14Accuracy (%) on the diagnosis-decision questions (complete a0

Interactive version: theaggregate.ai/model?slug=llama-3-8b-ultramedical · How It Works · Data refreshed daily, snapshot 2026-10-09.