ThaiSafetyBench - Discrimination, Exclusion, Toxicity, Hateful, Offensive: leaderboard

Metric: Attack Success Rate (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1GPT-52.19
2Claude Sonnet 4.53.49
3Qwen 2.5 72B Instruct6.27
4openthaigpt1.5-72B-instruct6.58
5Qwen 2.5 7B Instruct6.97
6GPT-4o7.87
7openthaigpt1.5-7B-instruct8.47
8Llama 3.3 70B Instruct8.67
9Gemma 3 12B (IT)8.86
10SeaLLMs-v3-7B-Chat8.96
11llama3.1-typhoon2-70B-instruct10.26
12Llama 3.1 70B Instruct15.64
13Gemma 3 4B (IT)18.55
14Llama 3.2 3B Instruct24.32
15Llama 3.1 8B Instruct25.5

Interactive version: theaggregate.ai/benchmark?slug=thaisafetybench-discrimination-exclusion-toxicity-hateful-offensive · How It Works · Data refreshed daily, snapshot 2026-09-19.