BLUFF - Veracity (Native Prompt, Big-Head): leaderboard
Metric: Macro-F1 (%) of zero-shot binary real/fake news classification (macro-F1, %) on the BLUFF test split, the decoder LLM answering with temperature 0.1 and at most 10 new tokens, with the prompt in the target language (native prompt), averaged over the 16 big-head (high-resource) languages; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Llama 3.2 1B Instruct | 51.3 | #1518 |
| 2 | Gemma 3 1B (IT) | 47 | #1484 |
| 3 | Qwen 3 0.6B | 45.9 | #1440 |
| 4 | Gemma 3 270M (IT) | 41.1 | #1562 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=bluff-veracity-native-prompt-big-head · How It Works · Data refreshed daily, snapshot 2026-10-11.