LexNeo-Bench (KG-Flat Context): leaderboard

Metric: Accuracy (%; borrowing-type classification of a highlighted token in a Luxembourgish RTL news sentence among native, French loan and German loan (English loan offered as a distractor label); 1,000 items per class; outputs that cannot be mapped to a label are dropped; temperature 0; a fixed list of about 20 productive LuxBorrow adaptation patterns prepended to every prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 3 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct36.3
2Gemma 3 27B (IT)33.7
3Gemma 3 12B (IT)30.3

Interactive version: theaggregate.ai/benchmark?slug=lexneo-bench-kg-flat-context · How It Works · Data refreshed daily, snapshot 2026-09-26.