SommBench - Wine Feature Completion (Danish): leaderboard

Metric: Wine feature completion accuracy (%) with Danish output: attribute accuracy on SommBench's 1,000 wine profiles with one to three masked attributes (type, country, region, grapes, dryness, body, acidity, sugar, alcohol) completed as structured JSON in the target language, numeric attributes counted correct within 5% error, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o64#333
2Gemini 2.5 Pro63#145
3Gemini 2.5 Flash63#237
4GPT-4.163#240
5GPT-4.1 Mini62#346
6Grok 4 Fast62#242
7GPT-4o Mini61#588
8Grok 461#169
9GPT-558#91
10Gemini 2.5 Flash Lite56#413
11GPT-4.1 Nano49#716
12GPT-OSS-120B (Low)39#330 (GPT-OSS-120B)
13GPT-OSS-120B (Medium)35#330 (GPT-OSS-120B)
14GPT-OSS-20B (Medium)34#499 (GPT-OSS-20B)
15GPT-OSS-20B (Low)32#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-feature-completion-danish · How It Works · Data refreshed daily, snapshot 2026-10-11.