Llama 3.1 70B Cpt Sea Lionv3 Instruct: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1594 ± 1, rank #278 of 1392 rated models, from 52 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SeaEval - Emotion - SST-2 (Zero-Shot)95.3Accuracy (%)97.9
SeaEval - Emotion - IndoEmotion (Zero-Shot)73.86Accuracy (%)95.8
SeaEval - Fundamental NLP Tasks - C3 (Zero-Shot)96.41Accuracy (%)95.8
SeaEval - Fundamental NLP Tasks - QNLI (Zero-Shot)92Accuracy (%)95.7
Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU)37.53SacreBLEU95.5
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++)55.15chrF++94
Thai LLM MC - m3exam_tha_seacrowd_qa63.1Accuracy (%)92.8
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU)28.92BLEU92.5
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU)39.87SacreBLEU92.5
Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE1)71.63ROUGE192.5
Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL)71.45ROUGEL92.5
SeaEval - Cultural Reasoning - US-Eval (Zero-Shot)86.92Accuracy (%)91.7

Interactive version: theaggregate.ai/model?slug=llama-3-1-70b-cpt-sea-lionv3-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.