FACTS Leaderboard — leaderboard

FACTS grounding and quality evaluation: measures both factual grounding accuracy and response quality, combining them into a single score for open-source LLMs.

Metric: Combined Score (%). Source: huggingface.co. Status: saturation imminent. 34 models tracked.

Top models

#ModelScore
1DeepSeek R1 Distill Qwen 14B45.76
2Llama 3.3 70B Instruct42.55
3Qwen 3 30B A3B42.55
4Qwen 3 4B42.55
5Qwen 3 32B41.7
6DeepSeek R1 0528 Qwen3 8B41.1
7DeepSeek R1 Distill Llama 8B40.68
8Qwen 3 8B40
9Qwen 3 14B38.3
10Gemma 3 27B (IT)37.8
11Qwen 2.5 VL 32B Instruct35.74
12Llama 3.1 70B Instruct33.47
13Gemma 3 12B (IT)31.3
14Gemma 3 4B (IT)30
15Qwen 3 1.7B29.79

Interactive version: theaggregate.ai/benchmark?slug=facts-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.