LiveFact: leaderboard

Metric: Average (%) of the twelve LiveFact November 2025 scores (accuracy and macro-F1 in Classification and Inference modes at evidence offsets of -3, 0 and +3 days), 4,392 claims from 737 news events, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 18 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B 2507 Instruct72.4
2GPT-OSS-120B72.13
3GPT-5.172.02
4GPT-5.271.52
5Qwen 3 30B A3B 2507 Instruct69.46
6GPT-4o (2024-08-06)67.11
7GPT-4o Mini (2024-07-18)66.12
8Qwen 3 32B66.03
9Qwen 3 8B63.62
10DeepSeek V3.161.48
11GPT-OSS-20B57.05
12Llama 3.3 70B Instruct55.85
13Qwen 3 4B 2507 Instruct55.46
14Kimi K2 (Thinking)54.6
15Llama 3.1 8B Instruct49.96

Interactive version: theaggregate.ai/benchmark?slug=livefact · How It Works · Data refreshed daily, snapshot 2026-10-07.