MultiCW (Guided Answer): leaderboard
Metric: Accuracy (%) of zero-shot binary check-worthiness classification (claim worth fact-checking or not) on the 18,444-sample MultiCW test split, balanced across 16 languages, two writing styles (noisy social media and structured news) and the two classes, with the short Guided Answer prompt, decisions extracted from the response with a strict yes/no fallback prompt; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-4o Mini | 76 | #588 |
| 2 | GPT-4.1 | 73 | #240 |
| 3 | GPT-4o | 72 | #333 |
| 4 | GPT-4.1 Mini | 68 | #346 |
| 5 | Llama 3.3 70B Instruct | 67 | #520 |
| 6 | Mistral Nemo 12B | 65 | #1239 |
| 7 | Claude 3.7 Sonnet | 61 | #241 |
| 8 | Claude 3.5 Haiku | 59 | #553 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=multicw-guided-answer · How It Works · Data refreshed daily, snapshot 2026-10-11.