Sailor-7B-Chat: benchmark results
Provider: Sea AI Lab. Released 2024-03-02. Access: Open.
Unified ELO 1425 ± 1, rank #1170 of 1553 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Thai LLM NLU - wisesight_thai_sentiment_seacrowd_text | 45.53 | Accuracy (%) | 60.9 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (chrF++) | 57.38 | chrF++ | 53.7 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU) | 32.53 | SacreBLEU | 52.2 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU) | 29.86 | SacreBLEU | 49.3 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++) | 48.35 | chrF++ | 47.8 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (BLEU) | 45.87 | BLEU | 47.8 |
| Thai LLM NLU - xcopa_tha_seacrowd_qa | 79.6 | Accuracy (%) | 45.7 |
| Thai LLM NLU - xnli.tha_seacrowd_pairs | 33.49 | Accuracy (%) | 41.3 |
| Thai LLM - Knowledge III | 3.2 | Rating (0-10) | 38 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU) | 20.1 | BLEU | 37.3 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE2) | 24.24 | ROUGE2 | 34.3 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL) | 31.57 | ROUGEL | 34.3 |
Interactive version: theaggregate.ai/model?slug=sailor-7b-chat · How It Works · Data refreshed daily, snapshot 2026-09-08.