MultiCW (Guided Answer): leaderboard

Metric: Accuracy (%) of zero-shot binary check-worthiness classification (claim worth fact-checking or not) on the 18,444-sample MultiCW test split, balanced across 16 languages, two writing styles (noisy social media and structured news) and the two classes, with the short Guided Answer prompt, decisions extracted from the response with a strict yes/no fallback prompt; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o Mini76#588
2GPT-4.173#240
3GPT-4o72#333
4GPT-4.1 Mini68#346
5Llama 3.3 70B Instruct67#520
6Mistral Nemo 12B65#1239
7Claude 3.7 Sonnet61#241
8Claude 3.5 Haiku59#553

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=multicw-guided-answer · How It Works · Data refreshed daily, snapshot 2026-10-11.