Llama 3 8B Cpt Sea Lionv2 Base: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1410 ± 21, rank #1121 of 1605 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SeaEval - Cultural Reasoning - PH-Eval (Few-Shot) | 54 | Accuracy (%) | 66.7 |
| SeaEval - Emotion - IndoEmotion (Few-Shot) | 57.27 | Accuracy (%) | 66.7 |
| SeaEval - Multilingual Reasoning - IndoMMLU (Few-Shot) | 50.42 | Accuracy (%) | 66.7 |
| SeaEval - Cultural Reasoning - SG-Eval v1 Cleaned (Few-Shot) | 64.71 | Accuracy (%) | 50 |
| SeaEval - Fundamental NLP Tasks - C3 (Few-Shot) | 79.96 | Accuracy (%) | 50 |
| Sahabat-AI - Indonesian Average | 38.5 | Average Score (%) | 43.8 |
| Sahabat-AI - Javanese (JV) | 38.94 | Score (%) | 43.8 |
| Sahabat-AI - Sundanese (SU) | 34.11 | Score (%) | 43.8 |
| SEA LLM Leaderboard - SeaExam - Thai (Private) | 43.4 | Private Accuracy (%) | 40 |
| SEA LLM Leaderboard - SeaExam - Thai (Public) | 51.6 | Public Accuracy (%) | 40 |
| SEA LLM Leaderboard - SeaExam - Indonesian (Public) | 49.2 | Public Accuracy (%) | 37.2 |
| SEA LLM Leaderboard - SeaExam | 52.2 | Private Average Score (%) | 36.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-8b-cpt-sea-lionv2-base · How It Works · Data refreshed daily, snapshot 2026-09-26.