Compl-AI Board — leaderboard
EU AI Act compliance evaluation: measures LLMs across 27 dimensions including bias (BBQ), toxicity, robustness, calibration, and safety using the Compl-AI framework.
Metric: Average Compliance Score (%). Source: huggingface.co. Status: saturated. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4 Preview (1106) | 86.44 |
| 2 | Claude 3 Opus | 84.9 |
| 3 | Gemini 1.5 Flash (001) | 80.44 |
| 4 | GPT-3.5 Turbo (0125) | 76.99 |
| 5 | Yi 34B (Chat) | 71.94 |
| 6 | Qwen 1.5 72B Chat | 71.52 |
| 7 | Bielik-11B-v2.3-Instruct | 71.12 |
| 8 | Llama 2 70B Chat (HF) | 70.1 |
| 9 | Mixtral 8x7B Instruct (v0.1) | 69.91 |
| 10 | Mistral 7B Instruct (v0.3) | 67.96 |
| 11 | Mistral 7B Instruct (v0.2) | 67.1 |
| 12 | Llama 2 13B Chat Base | 66.13 |
| 13 | Mistral-7B-v0.3 | 65.67 |
| 14 | Llama 2 7B Chat | 63.09 |
| 15 | Gemma 2 9B | 57.99 |
Interactive version: theaggregate.ai/benchmark?slug=compl-ai-board · How the rankings work · Data refreshed daily, snapshot 2026-07-22.