SLM Trustworthiness - Robustness: leaderboard

Metric: TrustLLM robustness accuracy (%; sub-task mean). Source: arxiv.org. 11 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct71.62
2Qwen 2.5 1.5B Instruct69.38
3Gemma 1.1 7B (IT)68.9
4Qwen 2.5 3B Instruct68.79
5Qwen 2.5 7B Instruct67.05
6Llama 3.2 3B Instruct66.42
7Qwen 2.5 0.5B Instruct62.42
8SmolLM2-360M-Instruct60.16
9Llama 3.2 1B Instruct59.43
10h2o-danube3-500m-chat56.85

Interactive version: theaggregate.ai/benchmark?slug=slm-trustworthiness-robustness · How It Works · Data refreshed daily, snapshot 2026-09-19.