HardMTBench - English to Chinese: leaderboard

Metric: GEMBA-DA direct assessment score (0-100) of the English-to-Chinese translations, judged by gpt-oss-120b, on HardMTBench, 10,000 domain-balanced, difficulty-selected Chinese-English pairs from 12 knowledge-intensive domains, each translated in both directions under one unified protocol; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)92.78
2GPT-5.592.66
3GPT-5.5 (Non-reasoning)92.6
4Qwen 3.5 397B A17B92.13
5Qwen 3.5 397B A17B (Non-reasoning)91.7
6DeepSeek V4 Pro (Non-reasoning)91.33
7Qwen 3.6 35B A3B91.05
8Hy-MT2-30B-A3B89.71
9Qwen 3.6 35B A3B (Non-reasoning)89.13
10Gemma 4 31B (IT) (Thinking)89.03
11gemma-4-26B-A4B-it (Thinking)88.53
12Gemma 4 31B (IT)88.25
13Gemma 4 26B A4B (IT)87.57
14Gemma 4 E4B (IT) (Thinking)83.24
15gemma-4-E4B-it82.82

Interactive version: theaggregate.ai/benchmark?slug=hardmtbench-english-to-chinese · How It Works · Data refreshed daily, snapshot 2026-10-07.