ClinicBench — leaderboard

Clinical NLP benchmark: evaluates LLMs across multiple clinical tasks including diagnosis, treatment planning, and medical document understanding.

Metric: Average Score. Source: huggingface.co. Status: years away from saturation. 23 models tracked.

Top models

#ModelScore
1GPT-456.96
2Llama 3 70B52.7
3Claude 246.35
4GPT-3.5 Turbo45.06
5Llama 2 70B39.7
6vicuna-13B36.48
7Llama 2 13B36.32
8Vicuna-7B34.61
9Llama 2 7B33.68

Interactive version: theaggregate.ai/benchmark?slug=clinicbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.