LexNeo-Bench (KG-Flat Context): leaderboard
Metric: Accuracy (%; borrowing-type classification of a highlighted token in a Luxembourgish RTL news sentence among native, French loan and German loan (English loan offered as a distractor label); 1,000 items per class; outputs that cannot be mapped to a label are dropped; temperature 0; a fixed list of about 20 productive LuxBorrow adaptation patterns prepended to every prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 36.3 |
| 2 | Gemma 3 27B (IT) | 33.7 |
| 3 | Gemma 3 12B (IT) | 30.3 |
Interactive version: theaggregate.ai/benchmark?slug=lexneo-bench-kg-flat-context · How It Works · Data refreshed daily, snapshot 2026-09-26.