CanMT - Chinese to English: leaderboard

Metric: Translation score (1-7) on the 125 Chinese-to-English CanMT sentences containing culture-specific items, GPT-5-nano judge with the reference translation scoring contextual accuracy, cultural adaptation, functional equivalence, fidelity and naturalness on 7-point Likert scales, averaged per sentence; default translation prompt; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1GPT-45.52
2Grok 4.15.43
3DeepSeek V3.25.41
4Gemini 2.5 Flash Lite5.4
5Qwen 2.5 72B Instruct5.37
6GPT-4o5.35
7Qwen 3 14B (Non-reasoning)5.31
8DeepSeek R15.28
9Qwen 2.5 14B Instruct5.26
10Qwen 2.5 32B Instruct5.22
11Qwen 3 32B (Non-reasoning)5.19
12Llama 3.3 70B Instruct5.14
13Qwen 3 8B (Non-reasoning)5.03
14Qwen 2.5 7B Instruct4.97
15Qwen 3 4B (Non-reasoning)4.83

Interactive version: theaggregate.ai/benchmark?slug=canmt-chinese-to-english · How It Works · Data refreshed daily, snapshot 2026-10-07.