LiveFact - Inference Macro-F1: leaderboard

Metric: Macro-F1 (%) in Inference Mode (label adjusted to what the evidence slice supports) with the evidence available three days after the event, LiveFact November 2025; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1GPT-5.181.99
2Qwen 3 235B A22B 2507 Instruct81.94
3GPT-OSS-120B81.31
4GPT-5.279.57
5Qwen 3 30B A3B 2507 Instruct76.4
6GPT-4o (2024-08-06)72.98
7Qwen 3 32B69.91
8GPT-4o Mini (2024-07-18)69.45
9Qwen 3 8B68.3
10DeepSeek V3.166.38
11Qwen 3 4B 2507 Instruct64.27
12Llama 3.3 70B Instruct63.93
13GPT-OSS-20B62.34
14Kimi K2 (Thinking)59.15
15Llama 3.1 8B Instruct45.14

Interactive version: theaggregate.ai/benchmark?slug=livefact-inference-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.