llama3.1-typhoon2-8B-instruct: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1480 ± 27, rank #1280 of 2656 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE1)76.61ROUGE197
Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGE2)63.19ROUGE297
Thai LLM NLG - iapp_squad_seacrowd_qa (ROUGEL)76.42ROUGEL97
Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (BLEU)50.56BLEU85.1
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU)27.36BLEU83.6
Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGE1)30.3ROUGE182.1
Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGE2)10.56ROUGE282.1
Thai LLM NLG - xl_sum_tha_seacrowd_t2t (ROUGEL)21.02ROUGEL80.6
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU)37.26SacreBLEU77.6
Thai LLM NLG - flores200_tha_Thai_eng_Latn_seacrowd_t2t (SacreBLEU)33.92SacreBLEU77.6
Thai LLM - Knowledge III4.55Rating (0-10)74.6
Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++)52.72chrF++74.6

Interactive version: theaggregate.ai/model?slug=llama3-1-typhoon2-8b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.