SeaLLMs-v3-7B-Chat — benchmark results

Alibaba DAMO Academy's Southeast Asian chat model built on Qwen2-7B, covering 12 regional languages from Indonesian to Burmese (July 2024). Provider: Sea AI Lab. Released 2024-07-03. Access: Open.

Unified ELO 1465 ± 7, rank #930 of 1776 rated models, from 166 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task81.06Accuracy (%)89.5
SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot)55Accuracy (%)89.1
Open Arabic LLM - Alghafa Multiple Choice Sentiment Task42.44Accuracy (%)88.6
SeaEval - Multilingual Reasoning - C-Eval (Zero-Shot)76.59Accuracy (%)82.6
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task54.5Accuracy (%)79.9
SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot)81.9Accuracy (%)79.2
SeaEval - Cultural Reasoning - SG-Eval (Zero-Shot)71.84Accuracy (%)76.2
SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot)76.84Accuracy (%)76.2
SeaEval - Dialogue - SAMSum (Zero-Shot)29.6Average ROUGE (0-100)75
SeaEval - Emotion - SST-2 (Zero-Shot)94.04Accuracy (%)75
Open Arabic LLM - Arabic MMLU HT Machine Learning43.75Accuracy (%)71.6
Open Arabic LLM - Madinah QA Arabic Language (Grammar)42.74Accuracy (%)70.4

Interactive version: theaggregate.ai/model?slug=seallms-v3-7b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.