Translation (Lechmazur) — leaderboard
Round-trip translation quality benchmark: English to target language and back across 10 languages, scoring preservation of meaning and fluency.
Metric: Mean Score. Source: github.com. Status: saturation imminent. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 (Medium) | 8.69 |
| 2 | Grok 4 | 8.57 |
| 3 | Claude Opus 4.1 (Non-reasoning) | 8.56 |
| 4 | Gemini 2.5 Pro | 8.53 |
| 5 | Qwen 3 Max (Preview) | 8.32 |
| 6 | Mistral Medium 3.1 | 8.29 |
| 7 | Kimi K2 0905 | 8.29 |
Interactive version: theaggregate.ai/benchmark?slug=translation-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.