MedS-Bench - Concept Explanation — leaderboard
Metric: Score. Source: henrychur.github.io. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 | 19.37 |
| 2 | Llama3.1-Aloe-Beta-8B | 13.63 |
| 3 | Claude 3.5 Sonnet (20240620) | 12.56 |
Interactive version: theaggregate.ai/benchmark?slug=meds-bench-concept-explanation · How the rankings work · Data refreshed daily, snapshot 2026-07-22.