Compl-AI Board — leaderboard

EU AI Act compliance evaluation: measures LLMs across 27 dimensions including bias (BBQ), toxicity, robustness, calibration, and safety using the Compl-AI framework.

Metric: Average Compliance Score (%). Source: huggingface.co. Status: saturated. 15 models tracked.

Top models

#ModelScore
1GPT-4 Preview (1106)86.44
2Claude 3 Opus84.9
3Gemini 1.5 Flash (001)80.44
4GPT-3.5 Turbo (0125)76.99
5Yi 34B (Chat)71.94
6Qwen 1.5 72B Chat71.52
7Bielik-11B-v2.3-Instruct71.12
8Llama 2 70B Chat (HF)70.1
9Mixtral 8x7B Instruct (v0.1)69.91
10Mistral 7B Instruct (v0.3)67.96
11Mistral 7B Instruct (v0.2)67.1
12Llama 2 13B Chat Base66.13
13Mistral-7B-v0.365.67
14Llama 2 7B Chat63.09
15Gemma 2 9B57.99

Interactive version: theaggregate.ai/benchmark?slug=compl-ai-board · How the rankings work · Data refreshed daily, snapshot 2026-07-22.