Hy-MT2-1.8B: benchmark results

Provider: Other.

Unified ELO 1448 ± 63, rank #999 of 1629 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
IFMTBench - Single-Constraint76.76IF_Score (%): hard constraints checked by deterministic veri50
IFMTBench - Multi-Constraint Translation Quality67.8xCOMET-XXL translation quality (%) on the 2,838 multi-constr42.9
IFMTBench69.36IF$_\text{T}$ (self-reported)28.6
IFMTBench - Multi-Constraint57.61IF_Score (%): hard constraints checked by deterministic veri28.6
IFMTBench - Single-Constraint Translation Quality80.74xCOMET-XXL translation quality (%) on the 4,506 single-const28.6
LM Market Cap LMC Score40LMC Score (0-100)28.4
HardMTBench - Chinese to English xCOMET61.94xCOMET-XXL score (%) of the Chinese-to-English translations,23.8
HardMTBench - English to Chinese xCOMET70.72xCOMET-XXL score (%) of the English-to-Chinese translations,23.8
HardMTBench - English to Chinese82.28GEMBA-DA direct assessment score (0-100) of the English-to-C14.3
HardMTBench78.22HardMTBench zh-en GEMBA-DA (self-reported)9.5
Tinybird AI SQL Benchmark - Success Rate42Questions answered with a valid query within 3 retries (%)3.4
Tinybird AI SQL Benchmark - First-Attempt Success Rate20Questions answered with a valid query on the first attempt (3

Interactive version: theaggregate.ai/model?slug=hy-mt2-1-8b · How It Works · Data refreshed daily, snapshot 2026-10-07.