MedES - Ethical Reasoning: leaderboard

Metric: Comprehensive score (-1 to 1; 4,004 open-ended medical-ethics queries: -1 per risky response, otherwise the mean of four binary quality checks, fine-tuned QwQ-32B evaluator). Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1DeepSeek R10.81
2DeepSeek V30.76
3GPT-4 Turbo0.44
4GPT-40.34
5DeepSeek-R1-Distill-Qwen-7B0.23
6GPT-3.50.22

Interactive version: theaggregate.ai/benchmark?slug=medes-ethical-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-25.