MultiCW (CoT): leaderboard

Metric: Accuracy (%) of zero-shot binary check-worthiness classification (claim worth fact-checking or not) on the 18,444-sample MultiCW test split, balanced across 16 languages, two writing styles (noisy social media and structured news) and the two classes, with a chain-of-thought prompt built from the CLEF annotation guidelines (CoT_CLEF_on_Q), decisions extracted from the response with a strict yes/no fallback prompt; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScoreOverall rank
1Claude 3.5 Haiku75#553
2Claude 3.7 Sonnet72#241
3Llama 3.3 70B Instruct71#520
4GPT-4o69#333
5GPT-4o Mini69#588
6Mistral Nemo 12B66#1239
7GPT-4.1 Mini65#346
8GPT-4.165#240

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=multicw-cot · How It Works · Data refreshed daily, snapshot 2026-10-11.