Thai LLM NLU - wisesight_thai_sentiment_seacrowd_text: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 70 models tracked.

Top models

#ModelScore
1Qwen 2 72B Instruct60.09
2DeepSeek R1 Distill Llama 70B56.76
3GPT-4o (2024-05-13)56.76
4Qwen 2.5 7B Instruct52.86
5Qwen 2.5 14B Instruct52.68
6GPT-4o Mini (2024-07-18)52.3
7Llama 3.3 70B Instruct51.82
8Tsunami-1.0-14B-Instruct51.18
9Llama 3.1 Nemotron 70B Instruct51.03
10Llama 3.1 8B Cpt Sea Lionv3 Instruct49.79
11Llama 3.1 70B Cpt Sea Lionv3 Instruct49.72
12Claude 3.5 Sonnet (20240620)49.61
13Llama 3.1 70B Instruct49.53
14Tsunami-0.5-7B-Instruct49.12
15Gemma 2 2B (IT)47.66

Interactive version: theaggregate.ai/benchmark?slug=thai-llm-nlu-wisesight-thai-sentiment-seacrowd-text · How It Works · Data refreshed daily, snapshot 2026-09-05.