ThaiSafetyBench - Misinformation Harms: leaderboard

Metric: Attack Success Rate (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1GPT-57.84
2SeaLLMs-v3-7B-Chat10.63
3Qwen 2.5 72B Instruct12.54
4Claude Sonnet 4.513.41
5Qwen 2.5 7B Instruct17.25
6openthaigpt1.5-72B-instruct17.94
7Llama 3.3 70B Instruct18.12
8llama3.1-typhoon2-70B-instruct18.82
9Llama 3.2 3B Instruct21.43
10openthaigpt1.5-7B-instruct23.52
11Llama 3.1 8B Instruct27.18
12llama3.1-typhoon2-8B-instruct29.09
13llama3.2-typhoon2-3B-instruct29.97
14Llama 3.1 70B Instruct30.31
15Gemma 3 12B (IT)30.49

Interactive version: theaggregate.ai/benchmark?slug=thaisafetybench-misinformation-harms · How It Works · Data refreshed daily, snapshot 2026-09-19.