BLUFF - Veracity (Native Prompt, Long-Tail): leaderboard

Metric: Macro-F1 (%) of zero-shot binary real/fake news classification (macro-F1, %) on the BLUFF test split, the decoder LLM answering with temperature 0.1 and at most 10 new tokens, with the prompt in the target language (native prompt), averaged over the 57 long-tail (low-resource) languages; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScoreOverall rank
1Llama 3.2 1B Instruct56.6#1518
2Gemma 3 1B (IT)49.8#1484
3Qwen 3 0.6B42.7#1440
4Gemma 3 270M (IT)42.4#1562

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=bluff-veracity-native-prompt-long-tail · How It Works · Data refreshed daily, snapshot 2026-10-11.