HardMTBench - English to Chinese xCOMET: leaderboard

Metric: xCOMET-XXL score (%) of the English-to-Chinese translations, a reference-based learned metric (0-1) times 100, on HardMTBench, 10,000 domain-balanced, difficulty-selected Chinese-English pairs from 12 knowledge-intensive domains, each translated in both directions under one unified protocol; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 22 models tracked.

Top models

#ModelScore
1Hy-MT2-30B-A3B75.87
2Gemini 3.1 Pro (Preview)74.15
3GPT-5.573.44
4GPT-5.5 (Non-reasoning)73.33
5Qwen 3.6 35B A3B72.57
6Qwen 3.5 397B A17B72.36
7Qwen 3.5 397B A17B (Non-reasoning)72.26
8Gemma 4 31B (IT) (Thinking)72.1
9gemma-4-26B-A4B-it (Thinking)71.81
10Qwen 3.6 35B A3B (Non-reasoning)71.8
11Gemma 4 26B A4B (IT)71.09
12Gemma 4 31B (IT)71.04
13DeepSeek V4 Pro (Non-reasoning)70.84
14Hy-MT2-1.8B70.72
15Gemma 4 E4B (IT) (Thinking)67.35

Interactive version: theaggregate.ai/benchmark?slug=hardmtbench-english-to-chinese-xcomet · How It Works · Data refreshed daily, snapshot 2026-10-07.