EarlyDx: leaderboard
Metric: Primary-track F1 (%): micro-averaged F1 of the predicted free-text diagnoses against the evidence-supported ED-encounter diagnoses (matched by a semantic judge), 6,975 test encounters with evidence clipped at hospital admission; zero-shot prompting; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 40 |
| 2 | GLM-5.2 | 37 |
| 3 | MedGemma-4B | 32 |
| 4 | Claude Opus 4.8 | 28 |
| 5 | Nemotron 3 Ultra | 25 |
Interactive version: theaggregate.ai/benchmark?slug=earlydx · How It Works · Data refreshed daily, snapshot 2026-09-29.