French LLM Leaderboard - Average — leaderboard

French LLM leaderboard aggregating IFEval FR, GPQA FR, and BAC FR to compare open models and chatbots on French instruction-following, graduate science QA, and exam-style reasoning.

Metric: Average Score (%). Source: huggingface.co. Status: saturation imminent. 55 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Llama 70B60.02
2Mistral Large 2 (Nov) Instruct (2411)55.19
3Llama 3.3 70B Instruct54.72
4DeepSeek R1 Distill Qwen 32B53.39
5Qwen 2.5 72B Instruct53.2
6calme-3.2-instruct-78B49.89
7Qwen 2.5 14B Instruct48.9
8Chocolatine-2-14B-Instruct-v2.048.71
9Phi-448.21
10Virtuoso-Lite46.8
11Mistral Small 345.02
12Mixtral 8x22B Instruct (v0.1)44.73
13Phi-3.5-MoE-instruct40.16
14Phi-3-medium-128k-instruct36.76
15calme-3.3-instruct-3B35.59

Interactive version: theaggregate.ai/benchmark?slug=french-llm-leaderboard-average · How the rankings work · Data refreshed daily, snapshot 2026-07-22.