MERaLiON-2-10B: benchmark results
Provider: A*STAR. Access: Open.
Unified ELO 1552 ± 18, rank #592 of 1607 rated models, from 163 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SEA-HELM (Indonesian) - Summarization | 19.28 | Normalized Score | 100 |
| SEA-SpeechBench - Emotion Recognition (SEA Prompt) | 20.34 | Judge-based accuracy (%; closed nine-class emotion label fro | 92.9 |
| SEA-SpeechBench - Speech Translation (SEA Prompt) | 19.52 | BLEU (0-100): corpus BLEU of the English translation of Sout | 92.9 |
| SEA-SpeechBench - ASR | 0.3 | Word or character error rate (raw ratio, lower is better; WE | 92.3 |
| SEA-HELM (Indonesian) - NLG | 55.33 | Normalized Score | 86 |
| SEA-HELM (Malay) - Toxicity Detection | 14.61 | Normalized Score | 86 |
| SEA-SpeechBench - Speech Translation (English Prompt) | 17.75 | BLEU (0-100): corpus BLEU of the English translation of Sout | 85.7 |
| SEA-HELM (Filipino) - Summarization | 21.67 | Normalized Score | 80.7 |
| SEA-HELM (Malay) - Safety | 42.24 | Normalized Score | 80.7 |
| SEA-HELM (Vietnamese) - Summarization | 16.96 | Normalized Score | 78.9 |
| SEA-SpeechBench - Temporal Localization (60-120 s) | 6.4 | Span-overlap F1 (%; predict the start and end time at which | 77.8 |
| SEA-SpeechBench - Speaker Recognition (English Prompt) | 48.36 | Macro-F1 (%; whether two clips come from the same speaker, o | 75 |
Interactive version: theaggregate.ai/model?slug=meralion-2-10b · How It Works · Data refreshed daily, snapshot 2026-09-29.