gemma2-9B-cpt-sea-lionv3-instruct: benchmark results
Provider: Other. Access: Open.
Unified ELO 1538 ± 1, rank #493 of 1392 rated models, from 69 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Thai LLM NLU - xnli.tha_seacrowd_pairs | 47.54 | Accuracy (%) | 95.7 |
| Thai LLM - Writing | 7.8 | Rating (0-10) | 94.4 |
| Thai LLM - Roleplay | 7.7 | Rating (0-10) | 90.8 |
| SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot) | 55 | Accuracy (%) | 89.1 |
| SeaEval - Emotion - IndoEmotion (Zero-Shot) | 73.41 | Accuracy (%) | 87.5 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (SacreBLEU) | 38.66 | SacreBLEU | 86.6 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (chrF++) | 53.63 | chrF++ | 85.1 |
| SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot) | 77.94 | Accuracy (%) | 84.8 |
| Thai LLM - Social Science | 8.2 | Rating (0-10) | 81.7 |
| Thai LLM NLG - flores200_eng_Latn_tha_Thai_seacrowd_t2t (BLEU) | 26.73 | BLEU | 79.1 |
| Thai LLM NLU - belebele_tha_thai_seacrowd_qa | 85 | Accuracy (%) | 75.4 |
| SeaEval - FLORES Translation - Malay-to-English (Zero-Shot) | 40.59 | BLEU (0-100) | 73.9 |
Interactive version: theaggregate.ai/model?slug=gemma2-9b-cpt-sea-lionv3-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.