Llama 3.1 70B Cpt Sea Lionv3 Instruct: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1594 ± 1, rank #278 of 1392 rated models, from 52 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SeaEval - Emotion - SST-2 (Zero-Shot) | 95.3 | Accuracy (%) | 97.9 |
| SeaEval - Emotion - IndoEmotion (Zero-Shot) | 73.86 | Accuracy (%) | 95.8 |
| SeaEval - Fundamental NLP Tasks - C3 (Zero-Shot) | 96.41 | Accuracy (%) | 95.8 |
| SeaEval - Fundamental NLP Tasks - QNLI (Zero-Shot) | 92 | Accuracy (%) | 95.7 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU) | 37.53 | SacreBLEU | 95.5 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++) | 55.15 | chrF++ | 94 |
| Thai LLM MC - m3exam_tha_seacrowd_qa | 63.1 | Accuracy (%) | 92.8 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU) | 28.92 | BLEU | 92.5 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU) | 39.87 | SacreBLEU | 92.5 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE1) | 71.63 | ROUGE1 | 92.5 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL) | 71.45 | ROUGEL | 92.5 |
| SeaEval - Cultural Reasoning - US-Eval (Zero-Shot) | 86.92 | Accuracy (%) | 91.7 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-70b-cpt-sea-lionv3-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.