TailNLG (Long-Tail Entities): leaderboard

Metric: chrF++ (0-100) against the TailNLG references for triples about long-tail (rare) Wikidata entities, averaged over English, Spanish and Italian; zero-shot prompt in the target language, 3 samples at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 3 12B (IT)60.38#655
2Gemma 3 4B (IT)54.32#971
3Qwen 2.5 7B Instruct54.27#846
4Llama 3.1 8B Instruct53.81#1018
5Qwen 2.5 3B Instruct50.1#1138
6Llama 3.2 3B Instruct49.75#1321

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tailnlg-long-tail-entities · How It Works · Data refreshed daily, snapshot 2026-10-11.