ThaiSafetyBench - General Prompt Attacks: leaderboard

Metric: Attack Success Rate (%). Source: huggingface.co. 24 models tracked.

Top models

#ModelScore
1GPT-51.76
2Claude Sonnet 4.53.3
3openthaigpt1.5-72B-instruct3.67
4Qwen 2.5 72B Instruct3.83
5openthaigpt1.5-7B-instruct5.59
6SeaLLMs-v3-7B-Chat5.7
7Qwen 2.5 7B Instruct5.75
8GPT-4o5.96
9Llama 3.3 70B Instruct6.39
10Gemma 3 12B (IT)9.27
11llama3.1-typhoon2-70B-instruct10.06
12Llama 3.1 70B Instruct12.52
13Llama 3.1 8B Instruct15.34
14Gemma 3 4B (IT)17.65
15Llama 3.2 3B Instruct17.89

Interactive version: theaggregate.ai/benchmark?slug=thaisafetybench-general-prompt-attacks · How It Works · Data refreshed daily, snapshot 2026-09-19.