Sailor2-8B-Chat: benchmark results
Sea AI Lab's Southeast Asian chat model, block-expanded from Qwen2.5-7B to 8B and continually pretrained on ~500B tokens for 15 languages (December 2024). Provider: Sea AI Lab. Released 2024-12-03. Access: Open.
Unified ELO 1457 ± 1, rank #912 of 1392 rated models, from 62 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Thai LLM - Knowledge III | 5.35 | Rating (0-10) | 93 |
| SeaEval - Emotion - IndoEmotion (Zero-Shot) | 73.64 | Accuracy (%) | 91.7 |
| Thai LLM - Writing | 7.45 | Rating (0-10) | 87.3 |
| SeaEval - Emotion - SST-2 (Zero-Shot) | 94.61 | Accuracy (%) | 83.3 |
| SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot) | 54.36 | Accuracy (%) | 82.6 |
| SeaEval - Fundamental NLP Tasks - QQP (Zero-Shot) | 82.05 | Accuracy (%) | 82.6 |
| Thai LLM NLU - xcopa_tha_seacrowd_qa | 91.2 | Accuracy (%) | 81.2 |
| SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot) | 77.7 | Accuracy (%) | 78.3 |
| Thai LLM - Social Science | 8 | Rating (0-10) | 77.5 |
| Thai LLM - STEM | 6.65 | Rating (0-10) | 74.6 |
| IndoBias | 69.4 | Ideology and Religion IND (self-reported) | 74 |
| Thai LLM MC - m3exam_tha_seacrowd_qa | 55.21 | Accuracy (%) | 69.6 |
Interactive version: theaggregate.ai/model?slug=sailor2-8b-chat · How It Works · Data refreshed daily, snapshot 2026-09-05.