HardMTBench - Chinese to English xCOMET: leaderboard

Metric: xCOMET-XXL score (%) of the Chinese-to-English translations, a reference-based learned metric (0-1) times 100, on HardMTBench, 10,000 domain-balanced, difficulty-selected Chinese-English pairs from 12 knowledge-intensive domains, each translated in both directions under one unified protocol; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 22 models tracked.

Top models

#ModelScore
1Hy-MT2-30B-A3B69.43
2Qwen 3.5 397B A17B (Non-reasoning)65.77
3GPT-5.5 (Non-reasoning)65.76
4GPT-5.565.72
5Gemma 4 31B (IT) (Thinking)65.7
6Qwen 3.6 35B A3B65.62
7Qwen 3.6 35B A3B (Non-reasoning)65.44
8DeepSeek V4 Pro (Non-reasoning)65.09
9gemma-4-26B-A4B-it (Thinking)64.96
10Gemma 4 31B (IT)64.67
11Gemini 3.1 Pro (Preview)64.45
12Qwen 3.5 397B A17B64.3
13Gemma 4 26B A4B (IT)63.81
14Hy-MT2-1.8B61.94
15Gemma 4 E4B (IT) (Thinking)60.39

Interactive version: theaggregate.ai/benchmark?slug=hardmtbench-chinese-to-english-xcomet · How It Works · Data refreshed daily, snapshot 2026-10-07.