HELM AFR - MMLU Clinical Knowledge (Bambara): leaderboard

Metric: EM. Source: crfm.stanford.edu. 22 models tracked.

Top models

#ModelScore
1Gemini 2.0 Flash63.4
2Gemini 2.0 Flash Lite60
3Claude 3.7 Sonnet (20250219)56.98
4DeepSeek V347.17
5GPT-4o (2024-08-06)43.02
6Claude 3.5 Haiku (20241022)42.64
7Llama 3.1 70B Instruct40.75
8Qwen 2.5 72B Instruct40.75
9Llama 3.3 70B Instruct39.25
10GPT-3.5 Turbo (0125)35.85
11Llama 3.1 8B Instruct33.96
12GPT-4 Turbo32.83
13Gemma 2 27B (IT)30.94
14Mistral 7B Instruct (v0.3)27.17
15GPT-226.79

Interactive version: theaggregate.ai/benchmark?slug=helm-afr-mmlu-clinical-knowledge-bambara · How It Works · Data refreshed daily, snapshot 2026-09-08.