CARTE: leaderboard
Metric: Accuracy (%) on all 2,431 CARTE questions (question-weighted over the regional columns), 0-shot, five options per question (four answers and 'Je ne sais pas', which scores as wrong); questions generated by Gemini 3 Flash from human-selected French documents and filtered; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash | 91.9 |
| 2 | Gemma 3 12B (IT) | 74.7 |
| 3 | Mistral Nemo Instruct (2407) | 73.6 |
| 4 | Qwen 3.5 9B | 73.5 |
| 5 | Llama 3.1 8B Instruct | 71.7 |
| 6 | aya-expanse-8B | 67.4 |
| 7 | Mistral 7B Instruct (v0.3) | 66.6 |
| 8 | Qwen 3.5 4B | 65.8 |
| 9 | Mistral 7B Instruct (v0.2) | 61.1 |
| 10 | Llama 3.2 3B Instruct | 54.6 |
| 11 | Llama 3.2 1B Instruct | 47.2 |
| 12 | Luth-LFM2-1.2B | 45.2 |
| 13 | Lucie-7B | 40 |
| 14 | Mistral-7B-v0.1 | 27.9 |
| 15 | CroissantLLMChat-v0.1 | 20.1 |
Interactive version: theaggregate.ai/benchmark?slug=carte · How It Works · Data refreshed daily, snapshot 2026-09-29.