M-GATE - Grammatical Error Detection: leaderboard
Metric: Mean MCC over 30 languages (-1 to 1; temperature 0). Source: arxiv.org. 82 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 0.36 |
| 2 | Claude Fable 5 | 0.36 |
| 3 | Gemini 3.5 Flash | 0.3 |
| 4 | Gemini 3.5 Flash (High) | 0.28 |
| 5 | Gemini 3 Flash (Preview) | 0.23 |
| 6 | Grok 4.20 | 0.2 |
| 7 | Gemini 3.5 Flash (Minimal) | 0.19 |
| 8 | GPT-5.5 (High) | 0.18 |
| 9 | Gemini 3 Flash (Preview) (Minimal) | 0.18 |
| 10 | GPT-5.5 (Medium) | 0.16 |
| 11 | Claude Opus 4.6 (Thinking) | 0.15 |
| 12 | Gemini 2.5 Pro | 0.11 |
| 13 | Claude Sonnet 5 | 0.1 |
| 14 | Claude Opus 4.8 | 0.09 |
| 15 | Gemini 2.5 Flash | 0.08 |
Interactive version: theaggregate.ai/benchmark?slug=m-gate-grammatical-error-detection · How It Works · Data refreshed daily, snapshot 2026-09-19.