Llama 3 8B Cpt Sea Lionv2 Instruct: benchmark results
Provider: Meta. Access: Open.
Unified ELO 1451 ± 1, rank #939 of 1392 rated models, from 61 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SeaEval - Dialogue - SAMSum (Zero-Shot) | 30.7 | Average ROUGE (0-100) | 91.7 |
| SeaEval - Dialogue - DialogSum (Zero-Shot) | 25.78 | Average ROUGE (0-100) | 91.3 |
| Thai LLM - Knowledge III | 5.05 | Rating (0-10) | 88 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (BLEU) | 50.77 | BLEU | 86.6 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU) | 27.25 | BLEU | 80.6 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU) | 37.8 | SacreBLEU | 80.6 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU) | 34.14 | SacreBLEU | 80.6 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++) | 53.08 | chrF++ | 79.1 |
| SeaEval - Cultural Reasoning - SG-Eval v1 Cleaned (Zero-Shot) | 66.18 | Accuracy (%) | 73.8 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE1) | 54.04 | ROUGE1 | 71.6 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL) | 53.61 | ROUGEL | 71.6 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE2) | 42.74 | ROUGE2 | 70.1 |
Interactive version: theaggregate.ai/model?slug=llama-3-8b-cpt-sea-lionv2-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.