Hy-MT2-1.8B: benchmark results
Provider: Other.
Unified ELO 1448 ± 63, rank #999 of 1629 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| IFMTBench - Single-Constraint | 76.76 | IF_Score (%): hard constraints checked by deterministic veri | 50 |
| IFMTBench - Multi-Constraint Translation Quality | 67.8 | xCOMET-XXL translation quality (%) on the 2,838 multi-constr | 42.9 |
| IFMTBench | 69.36 | IF$_\text{T}$ (self-reported) | 28.6 |
| IFMTBench - Multi-Constraint | 57.61 | IF_Score (%): hard constraints checked by deterministic veri | 28.6 |
| IFMTBench - Single-Constraint Translation Quality | 80.74 | xCOMET-XXL translation quality (%) on the 4,506 single-const | 28.6 |
| LM Market Cap LMC Score | 40 | LMC Score (0-100) | 28.4 |
| HardMTBench - Chinese to English xCOMET | 61.94 | xCOMET-XXL score (%) of the Chinese-to-English translations, | 23.8 |
| HardMTBench - English to Chinese xCOMET | 70.72 | xCOMET-XXL score (%) of the English-to-Chinese translations, | 23.8 |
| HardMTBench - English to Chinese | 82.28 | GEMBA-DA direct assessment score (0-100) of the English-to-C | 14.3 |
| HardMTBench | 78.22 | HardMTBench zh-en GEMBA-DA (self-reported) | 9.5 |
| Tinybird AI SQL Benchmark - Success Rate | 42 | Questions answered with a valid query within 3 retries (%) | 3.4 |
| Tinybird AI SQL Benchmark - First-Attempt Success Rate | 20 | Questions answered with a valid query on the first attempt ( | 3 |
Interactive version: theaggregate.ai/model?slug=hy-mt2-1-8b · How It Works · Data refreshed daily, snapshot 2026-10-07.