CARTE-LV: leaderboard

Metric: Accuracy (%) on the 233 CARTE-LV questions on regional linguistic variation (grammar, discourse markers, pragmatics, register and region-specific lexical usage in 13 regions), 0-shot, five options including 'Je ne sais pas'; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 27 models tracked.

Top models

#ModelScore
1Gemini 3 Flash93.6
2Mistral Nemo Instruct (2407)74.2
3Gemma 3 12B (IT)68.7
4Qwen 3.5 9B65.7
5Llama 3.1 8B Instruct63.9
6Qwen 3.5 4B58.4
7aya-expanse-8B58.4
8Mistral 7B Instruct (v0.2)57.9
9Mistral 7B Instruct (v0.3)57.5
10Luth-LFM2-1.2B49.4
11Llama 3.2 3B Instruct48.9
12Lucie-7B44.2
13Llama 3.2 1B Instruct40.8
14gemma-4-E4B-it23.2
15Mistral-7B-v0.123.2

Interactive version: theaggregate.ai/benchmark?slug=carte-lv · How It Works · Data refreshed daily, snapshot 2026-09-29.