LLM Trustworthy - Stereotype — leaderboard

Metric: Trust Score (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1gemma-7B (IT)100
2Claude 2100
3GPT-4o (2024-05-13)99.67
4Llama 3 8B Instruct98.33
5Gemini 1.0 Pro98.33
6Llama 2 7B Chat97.6
7tulu-2-7B96.6
8zephyr-7B-beta92.6
9tulu-2-13B89.33
10GPT-4o Mini (2024-07-18)87.34
11falcon-7B Instruct87
12GPT-3.5 Turbo (0301)87
13mpt-7B-chat84.6
14vicuna-7B-v1.381
15Mistral-7B-OpenOrca79.33

Interactive version: theaggregate.ai/benchmark?slug=llm-trustworthy-stereotype · How the rankings work · Data refreshed daily, snapshot 2026-07-22.