Llama 3 8B Cpt Sea Lionv2.1 Instruct: benchmark results
Provider: Meta. Released 2024-08-01. Access: Open.
Unified ELO 1479 ± 1, rank #790 of 1392 rated models, from 67 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE1) | 74.58 | ROUGE1 | 95.5 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL) | 74.2 | ROUGEL | 95.5 |
| Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE2) | 59.15 | ROUGE2 | 94 |
| Thai LLM - Writing | 7.5 | Rating (0-10) | 90.1 |
| Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGE1) | 32.24 | ROUGE1 | 89.6 |
| Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGEL) | 21.93 | ROUGEL | 89.6 |
| Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGE2) | 11.33 | ROUGE2 | 88.1 |
| SeaEval - Dialogue - SAMSum (Zero-Shot) | 30.5 | Average ROUGE (0-100) | 87.5 |
| SeaEval - Dialogue - DialogSum (Zero-Shot) | 25.38 | Average ROUGE (0-100) | 78.3 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU) | 33.86 | SacreBLEU | 76.1 |
| Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (BLEU) | 49.59 | BLEU | 74.6 |
| SeaEval - Cultural Reasoning - SG-Eval v1 Cleaned (Zero-Shot) | 66.18 | Accuracy (%) | 73.8 |
Interactive version: theaggregate.ai/model?slug=llama-3-8b-cpt-sea-lionv2-1-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.