LiveFact (December 2025): leaderboard
Metric: Average (%) of the twelve LiveFact scores on the December 2025 release (4,222 claims), accuracy and macro-F1 in Classification and Inference modes at evidence offsets of -3, 0 and +3 days; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 235B A22B 2507 Instruct | 71.02 |
| 2 | Qwen 3 30B A3B 2507 Instruct | 68.32 |
| 3 | DeepSeek V3.1 | 66.76 |
| 4 | Qwen 3 8B | 61.15 |
| 5 | Qwen 3 4B 2507 Instruct | 55.81 |
| 6 | Llama 3.3 70B Instruct | 54.97 |
| 7 | Llama 3.1 8B Instruct | 49.87 |
Interactive version: theaggregate.ai/benchmark?slug=livefact-december-2025 · How It Works · Data refreshed daily, snapshot 2026-10-07.