LLM Trustworthy - Adversarial Demo: leaderboard

Metric: Trust Score (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1GPT-4o Mini (2024-07-18)88.49
2GPT-4o (2024-05-13)88.1
3GPT-3.5 Turbo (0301)81.28
4GPT-4 (0314)77.94
5Llama 3 8B Instruct75.54
6Gemini 1.0 Pro75.54
7Claude 272.97
8tulu-2-13B71.17
9zephyr-7B-beta68.68
10Mistral-7B-OpenOrca62.15
11tulu-2-7B60.49
12vicuna-7B-v1.357.99
13Llama 2 7B Chat55.54
14gemma-2B (IT)35.55
15falcon-7B Instruct33.95

Interactive version: theaggregate.ai/benchmark?slug=llm-trustworthy-adversarial-demo · How It Works · Data refreshed daily, snapshot 2026-09-05.