BLUFF - Veracity (Native Prompt, Long-Tail): leaderboard
Metric: Macro-F1 (%) of zero-shot binary real/fake news classification (macro-F1, %) on the BLUFF test split, the decoder LLM answering with temperature 0.1 and at most 10 new tokens, with the prompt in the target language (native prompt), averaged over the 57 long-tail (low-resource) languages; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Llama 3.2 1B Instruct | 56.6 | #1518 |
| 2 | Gemma 3 1B (IT) | 49.8 | #1484 |
| 3 | Qwen 3 0.6B | 42.7 | #1440 |
| 4 | Gemma 3 270M (IT) | 42.4 | #1562 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=bluff-veracity-native-prompt-long-tail · How It Works · Data refreshed daily, snapshot 2026-10-11.