TailNLG (Head Entities): leaderboard

Metric: chrF++ (0-100) against the TailNLG references for triples about head (popular) Wikidata entities, averaged over English, Spanish and Italian; zero-shot prompt in the target language, 3 samples at temperature 0.7; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScoreOverall rank
1Gemma 3 12B (IT)58.73#655
2Gemma 3 4B (IT)53.48#971
3Qwen 2.5 7B Instruct53.41#846
4Llama 3.1 8B Instruct53.34#1018
5Qwen 2.5 3B Instruct49.12#1138
6Llama 3.2 3B Instruct48.19#1321

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tailnlg-head-entities · How It Works · Data refreshed daily, snapshot 2026-10-11.