AI-Secure LLM Trustworthy Leaderboard: leaderboard

Average DecodingTrust score (UIUC AI-Secure lab, 2023) over eight perspectives: toxicity, stereotype bias, adversarial and OOD inputs, adversarial demonstrations, privacy, machine ethics, fairness.

Metric: Average Trust Score (%). Source: huggingface.co. Status: saturated. 35 models tracked.

Top models

#ModelScore
1Claude 284.52
2GPT-4o (2024-05-13)82.96
3Llama 3 8B Instruct80.61
4Gemini 1.0 Pro80.61
5GPT-4o Mini (2024-07-18)76.31
6Llama 2 7B Chat74.72
7GPT-3.5 Turbo (0301)72.45
8GPT-4 (0314)69.24
9gemma-2B (IT)67.18
10gemma-7B (IT)66.87
11tulu-2-13B66.51
12tulu-2-7B63.56
13zephyr-7B-beta63.24
14vicuna-7B-v1.360.62
15falcon-7B Instruct59.49

Interactive version: theaggregate.ai/benchmark?slug=ai-secure-llm-trustworthy-leaderboard · How It Works · Data refreshed daily, snapshot 2026-09-05.