Sailor2-8B-Chat — benchmark results
Sea AI Lab's Southeast Asian chat model, block-expanded from Qwen2.5-7B to 8B and continually pretrained on ~500B tokens for 15 languages (December 2024). Provider: Sea AI Lab. Released 2024-12-03. Access: Open.
Unified ELO 1456 ± 11, rank #967 of 1776 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SeaEval - Emotion - IndoEmotion (Zero-Shot) | 73.64 | Accuracy (%) | 91.7 |
| SeaEval - Emotion - SST-2 (Zero-Shot) | 94.61 | Accuracy (%) | 83.3 |
| SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot) | 54.36 | Accuracy (%) | 82.6 |
| SeaEval - Fundamental NLP Tasks - QQP (Zero-Shot) | 82.05 | Accuracy (%) | 82.6 |
| SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot) | 77.7 | Accuracy (%) | 78.3 |
| IndoBias | 69.4 | Ideology and Religion IND (self-reported) | 74 |
| SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot) | 71.43 | Accuracy (%) | 64.6 |
| SEA LLM Leaderboard - SeaExam | 64.2 | Private Average Score (%) | 62.2 |
| SeaEval - Fundamental NLP Tasks - RTE (Zero-Shot) | 81.23 | Accuracy (%) | 58.3 |
| SeaEval - Multilingual Reasoning - ZBench (Zero-Shot) | 51.52 | Accuracy (%) | 56.2 |
| SeaEval - Fundamental NLP Tasks - CoLA (Zero-Shot) | 79 | Accuracy (%) | 50 |
| SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot) | 64.17 | Accuracy (%) | 47.6 |
Interactive version: theaggregate.ai/model?slug=sailor2-8b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.