HELM WMT 2014 — leaderboard

Machine translation quality on WMT 2014 benchmark across multiple language pairs, scored by BLEU-4. Part of Stanford HELM.

Metric: BLEU-4 (%). Source: crfm.stanford.edu. Status: saturation imminent. 91 models tracked.

Top models

#ModelScore
1Palmyra X V3 (72B)26.19
2PaLM-2 (Unicorn)25.95
3PaLM-2 (Bison)24.08
4Claude 3 Opus (20240229)23.99
5Palmyra X V2 (33B)23.88
6Llama 3.1 405B Instruct23.81
7Gemini 1.5 Pro (002)23.13
8GPT-4o (2024-05-13)23.07
9Claude 3.5 Sonnet (20240620)22.9
10Nova Pro22.89
11Claude 3.5 Sonnet (20241022)22.56
12GPT-4o (2024-08-06)22.51
13Gemini 1.5 Flash (001)22.51
14Llama 3 70B22.46
15Llama 3.2 90B Vision Instruct22.41

Interactive version: theaggregate.ai/benchmark?slug=helm-wmt-2014 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.