Translation Set1->en COMET22: leaderboard

COMET-22 is a neural machine translation evaluation metric that uses an ensemble of two models: a COMET estimator trained with Direct Assessments and a multitask model that predicts sentence-level scores and word-level OK/BAD tags. It provides improved correlations with human judgments and increased robustness to critical errors compared to previous metrics.

Metric: Set1→en COMET22 (self-reported). Source: benchmarklist.com. 13 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro (002)89.1
2Claude 3.5 Sonnet89.1
3GPT-4o89
4Nova Pro89
5Gemini 1.5 Flash (002)88.8
6Nova Lite88.8
7GPT-4o Mini88.7
8Nova Micro88.7
9Gemini 1.5 Flash-8B (001)88.5
10Llama 3.2 90B88.5
11Claude 3.5 Haiku88.3
12Llama 3.1 8B86.5

Interactive version: theaggregate.ai/benchmark?slug=translation-set1-over-en-comet22 · How It Works · Data refreshed daily, snapshot 2026-09-05.