CARTE: leaderboard

Metric: Accuracy (%) on all 2,431 CARTE questions (question-weighted over the regional columns), 0-shot, five options per question (four answers and 'Je ne sais pas', which scores as wrong); questions generated by Gemini 3 Flash from human-selected French documents and filtered; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 27 models tracked.

Top models

#ModelScore
1Gemini 3 Flash91.9
2Gemma 3 12B (IT)74.7
3Mistral Nemo Instruct (2407)73.6
4Qwen 3.5 9B73.5
5Llama 3.1 8B Instruct71.7
6aya-expanse-8B67.4
7Mistral 7B Instruct (v0.3)66.6
8Qwen 3.5 4B65.8
9Mistral 7B Instruct (v0.2)61.1
10Llama 3.2 3B Instruct54.6
11Llama 3.2 1B Instruct47.2
12Luth-LFM2-1.2B45.2
13Lucie-7B40
14Mistral-7B-v0.127.9
15CroissantLLMChat-v0.120.1

Interactive version: theaggregate.ai/benchmark?slug=carte · How It Works · Data refreshed daily, snapshot 2026-09-29.