Sailor2-8B-Chat: benchmark results

Sea AI Lab's Southeast Asian chat model, block-expanded from Qwen2.5-7B to 8B and continually pretrained on ~500B tokens for 15 languages (December 2024). Provider: Sea AI Lab. Released 2024-12-03. Access: Open.

Unified ELO 1457 ± 1, rank #912 of 1392 rated models, from 62 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Thai LLM - Knowledge III5.35Rating (0-10)93
SeaEval - Emotion - IndoEmotion (Zero-Shot)73.64Accuracy (%)91.7
Thai LLM - Writing7.45Rating (0-10)87.3
SeaEval - Emotion - SST-2 (Zero-Shot)94.61Accuracy (%)83.3
SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot)54.36Accuracy (%)82.6
SeaEval - Fundamental NLP Tasks - QQP (Zero-Shot)82.05Accuracy (%)82.6
Thai LLM NLU - xcopa_tha_seacrowd_qa91.2Accuracy (%)81.2
SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot)77.7Accuracy (%)78.3
Thai LLM - Social Science8Rating (0-10)77.5
Thai LLM - STEM6.65Rating (0-10)74.6
IndoBias69.4Ideology and Religion IND (self-reported)74
Thai LLM MC - m3exam_tha_seacrowd_qa55.21Accuracy (%)69.6

Interactive version: theaggregate.ai/model?slug=sailor2-8b-chat · How It Works · Data refreshed daily, snapshot 2026-09-05.