SLM Trustworthiness - Machine Ethics: leaderboard

Metric: TrustLLM machine-ethics accuracy (%; sub-task mean). Source: arxiv.org. 11 models tracked.

Top models

#ModelScore
1Qwen 2.5 7B Instruct83.78
2Llama 3.2 3B Instruct82.88
3Gemma 1.1 7B (IT)82.19
4Llama 3.1 8B Instruct81.29
5Qwen 2.5 3B Instruct80.13
6Qwen 2.5 1.5B Instruct77.52
7Llama 3.2 1B Instruct73.98
8Qwen 2.5 0.5B Instruct65.71
9h2o-danube3-500m-chat61.92
10SmolLM2-360M-Instruct51.47

Interactive version: theaggregate.ai/benchmark?slug=slm-trustworthiness-machine-ethics · How It Works · Data refreshed daily, snapshot 2026-09-19.