LiveSecBench: leaderboard

Dynamic live safety benchmark for large language models across ethics, legality, privacy, factuality, and psychological health.

Metric: Overall Score (%). Source: livesecbench.intokentech.cn. Status: saturation imminent. 43 models tracked.

Top models

#ModelScore
1Claude Haiku 4.591.43
2Claude Sonnet 4.685.97
3GPT-5.284.72
4Qwen 3.5 397B A17B81.52
5Spark X279.18
6Kimi K2.574.79
7Doubao-Seed-1.670.83
8Qwen 3 235B A22B69.23
9MiniMax-M266.69
10GPT-OSS-120B66.63
11Seed 2.0 Pro63.04
12MiniMax-M2.561.65
13Gemini 3.1 Pro (Preview)58.16
14MiMo-V2-Flash57.23
15LongCat-Flash-Chat57.1

Interactive version: theaggregate.ai/benchmark?slug=livesecbench · How It Works · Data refreshed daily, snapshot 2026-09-05.