GaelEval - Linguistic Competence (English Prompt): leaderboard

Metric: Accuracy (%) on the 120-item Gaelic morphosyntactic MCQA with an English-language system prompt, exact string match, single call per item, zero-shot; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 19 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)80
2Gemini 3 Flash (Preview)77.5
3GPT-571.7
4Gemini 2.5 Flash62.5
5Claude Opus 4.650.8
6GPT-4o46.7
7Claude Haiku 4.543.3
8GPT-4.142.5
9DeepSeek R142.5
10GPT-5 Mini41.7
11GPT-5.240
12Llama 4 Maverick40
13GPT-4.1 Nano31.7
14GPT-5 Nano30.8
15GPT-4o Mini30

Interactive version: theaggregate.ai/benchmark?slug=gaeleval-linguistic-competence-english-prompt · How It Works · Data refreshed daily, snapshot 2026-10-07.