Sailor2-8B-Chat — benchmark results

Sea AI Lab's Southeast Asian chat model, block-expanded from Qwen2.5-7B to 8B and continually pretrained on ~500B tokens for 15 languages (December 2024). Provider: Sea AI Lab. Released 2024-12-03. Access: Open.

Unified ELO 1456 ± 11, rank #967 of 1776 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SeaEval - Emotion - IndoEmotion (Zero-Shot)73.64Accuracy (%)91.7
SeaEval - Emotion - SST-2 (Zero-Shot)94.61Accuracy (%)83.3
SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot)54.36Accuracy (%)82.6
SeaEval - Fundamental NLP Tasks - QQP (Zero-Shot)82.05Accuracy (%)82.6
SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot)77.7Accuracy (%)78.3
IndoBias69.4Ideology and Religion IND (self-reported)74
SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot)71.43Accuracy (%)64.6
SEA LLM Leaderboard - SeaExam64.2Private Average Score (%)62.2
SeaEval - Fundamental NLP Tasks - RTE (Zero-Shot)81.23Accuracy (%)58.3
SeaEval - Multilingual Reasoning - ZBench (Zero-Shot)51.52Accuracy (%)56.2
SeaEval - Fundamental NLP Tasks - CoLA (Zero-Shot)79Accuracy (%)50
SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot)64.17Accuracy (%)47.6

Interactive version: theaggregate.ai/model?slug=sailor2-8b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.