LiveFact - Classification Accuracy: leaderboard

Metric: Accuracy (%) in Classification Mode (Real, Fake or Ambiguous against the time-invariant label) with the evidence available three days after the event, LiveFact November 2025; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B 2507 Instruct82.08
2GPT-OSS-120B81.81
3GPT-5.181.01
4GPT-5.277.32
5Qwen 3 30B A3B 2507 Instruct77
6GPT-4o (2024-08-06)73.98
7Qwen 3 32B71.86
8GPT-4o Mini (2024-07-18)71.22
9Qwen 3 8B69.83
10Llama 3.3 70B Instruct69.76
11GPT-OSS-20B67.42
12Qwen 3 4B 2507 Instruct64.8
13DeepSeek V3.163.73
14Kimi K2 (Thinking)54.21
15Llama 3.1 8B Instruct50.2

Interactive version: theaggregate.ai/benchmark?slug=livefact-classification-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.