SommBench - Wine Feature Completion: leaderboard

Metric: Wine feature completion accuracy (%), mean over eight output languages: attribute accuracy on SommBench's 1,000 wine profiles with one to three masked attributes (type, country, region, grapes, dryness, body, acidity, sugar, alcohol) completed as structured JSON in the target language, numeric attributes counted correct within 5% error, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4o63#333
2Gemini 2.5 Flash63#237
3Gemini 2.5 Pro62#145
4GPT-4.162#240
5GPT-4o Mini62#588
6GPT-4.1 Mini61#346
7Grok 461#169
8Grok 4 Fast59#242
9GPT-557#91
10Gemini 2.5 Flash Lite57#413
11GPT-4.1 Nano52#716
12GPT-OSS-120B (Low)41#330 (GPT-OSS-120B)
13GPT-OSS-120B (Medium)40#330 (GPT-OSS-120B)
14GPT-OSS-20B (Medium)38#499 (GPT-OSS-20B)
15GPT-OSS-20B (Low)33#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-feature-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.