LLMEval-Med - Medical Safety and Ethics: leaderboard

Metric: Usability rate: share of responses a GPT-4o judge scores 4 or 5 of 5 against an expert checklist (%). Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1O1 Preview64.81
2O1 Mini63.3
3DeepSeek R159.63
4GPT-4o56.27
5DeepSeek V347.71

Interactive version: theaggregate.ai/benchmark?slug=llmeval-med-medical-safety-and-ethics · How It Works · Data refreshed daily, snapshot 2026-09-25.