SommBench - Wine Feature Completion (Swedish): leaderboard

Metric: Wine feature completion accuracy (%) with Swedish output: attribute accuracy on SommBench's 1,000 wine profiles with one to three masked attributes (type, country, region, grapes, dryness, body, acidity, sugar, alcohol) completed as structured JSON in the target language, numeric attributes counted correct within 5% error, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o64#333
2Gemini 2.5 Pro62#145
3GPT-4.162#240
4Gemini 2.5 Flash61#237
5GPT-4.1 Mini61#346
6GPT-4o Mini61#588
7GPT-559#91
8Grok 459#169
9Grok 4 Fast58#242
10Gemini 2.5 Flash Lite55#413
11GPT-4.1 Nano51#716
12GPT-OSS-120B (Low)40#330 (GPT-OSS-120B)
13GPT-OSS-120B (Medium)39#330 (GPT-OSS-120B)
14GPT-OSS-20B (Medium)38#499 (GPT-OSS-20B)
15GPT-OSS-20B (Low)35#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-feature-completion-swedish · How It Works · Data refreshed daily, snapshot 2026-10-11.