Med-CoReasoner - MMLU-ProX-Health - Spanish: leaderboard

Metric: Accuracy (%). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.2 (Low)75.69
2GPT-5.1 (Low)72.63
3GPT-4o70.6
4DeepSeek V3.267.89
5Qwen 2.5 72B Instruct61.51
6Qwen 2.5 32B Instruct60.41
7Llama 3.1 70B Instruct60.26
8Claude 3.5 Haiku59.97

Interactive version: theaggregate.ai/benchmark?slug=med-coreasoner-mmlu-prox-health-spanish · How It Works · Data refreshed daily, snapshot 2026-09-25.