Translation Set1->en COMET22 — leaderboard
COMET-22 is a neural machine translation evaluation metric that uses an ensemble of two models: a COMET estimator trained with Direct Assessments and a multitask model that predicts sentence-level scores and word-level OK/BAD tags. It provides improved correlations with human judgments and increased robustness to critical errors compared to previous metrics.
Source: www.amazon.science.
Interactive version: theaggregate.ai/benchmark?slug=translation-set1-over-en-comet22 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.