SeaLLMs-v3-7B-Chat — benchmark results
Alibaba DAMO Academy's Southeast Asian chat model built on Qwen2-7B, covering 12 regional languages from Indonesian to Burmese (July 2024). Provider: Sea AI Lab. Released 2024-07-03. Access: Open.
Unified ELO 1465 ± 7, rank #930 of 1776 rated models, from 166 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task | 81.06 | Accuracy (%) | 89.5 |
| SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot) | 55 | Accuracy (%) | 89.1 |
| Open Arabic LLM - Alghafa Multiple Choice Sentiment Task | 42.44 | Accuracy (%) | 88.6 |
| SeaEval - Multilingual Reasoning - C-Eval (Zero-Shot) | 76.59 | Accuracy (%) | 82.6 |
| Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task | 54.5 | Accuracy (%) | 79.9 |
| SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot) | 81.9 | Accuracy (%) | 79.2 |
| SeaEval - Cultural Reasoning - SG-Eval (Zero-Shot) | 71.84 | Accuracy (%) | 76.2 |
| SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot) | 76.84 | Accuracy (%) | 76.2 |
| SeaEval - Dialogue - SAMSum (Zero-Shot) | 29.6 | Average ROUGE (0-100) | 75 |
| SeaEval - Emotion - SST-2 (Zero-Shot) | 94.04 | Accuracy (%) | 75 |
| Open Arabic LLM - Arabic MMLU HT Machine Learning | 43.75 | Accuracy (%) | 71.6 |
| Open Arabic LLM - Madinah QA Arabic Language (Grammar) | 42.74 | Accuracy (%) | 70.4 |
Interactive version: theaggregate.ai/model?slug=seallms-v3-7b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.