SLM Trustworthiness - Fairness: leaderboard

Metric: TrustLLM fairness accuracy (%; sub-task mean). Source: arxiv.org. 11 models tracked.

Top models

#ModelScore
1Qwen 2.5 1.5B Instruct71.82
2Llama 3.2 1B Instruct70.81
3Qwen 2.5 0.5B Instruct48.99
4SmolLM2-360M-Instruct47.45
5Gemma 1.1 7B (IT)44.73
6Llama 3.2 3B Instruct43.3
7h2o-danube3-500m-chat43.27
8Llama 3.1 8B Instruct36.03
9Qwen 2.5 3B Instruct29.99
10Qwen 2.5 7B Instruct28.33

Interactive version: theaggregate.ai/benchmark?slug=slm-trustworthiness-fairness · How It Works · Data refreshed daily, snapshot 2026-09-19.