Med-CoReasoner - Global-MMLU-Medical - Swahili: leaderboard

Metric: Accuracy (%). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.2 (Low)81.86
2GPT-5.1 (Low)79.19
3GPT-4o76.08
4DeepSeek V3.268.11
5Llama 3.1 70B Instruct65.58
6Claude 3.5 Haiku56.65
7Qwen 2.5 32B Instruct51.76
8Qwen 2.5 72B Instruct50.9

Interactive version: theaggregate.ai/benchmark?slug=med-coreasoner-global-mmlu-medical-swahili · How It Works · Data refreshed daily, snapshot 2026-09-25.