SommBench - Wine Feature Completion (Italian): leaderboard

Metric: Wine feature completion accuracy (%) with Italian output: attribute accuracy on SommBench's 1,000 wine profiles with one to three masked attributes (type, country, region, grapes, dryness, body, acidity, sugar, alcohol) completed as structured JSON in the target language, numeric attributes counted correct within 5% error, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 21 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Flash65#237
2GPT-4o64#333
3Gemini 2.5 Pro62#145
4GPT-4.162#240
5GPT-4o Mini62#588
6Grok 460#169
7GPT-559#91
8GPT-4.1 Mini59#346
9Grok 4 Fast57#242
10GPT-4.1 Nano56#716
11Gemini 2.5 Flash Lite55#413
12GPT-OSS-120B (Medium)44#330 (GPT-OSS-120B)
13GPT-OSS-20B (Medium)43#499 (GPT-OSS-20B)
14GPT-OSS-120B (Low)43#330 (GPT-OSS-120B)
15GPT-OSS-20B (Low)33#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-feature-completion-italian · How It Works · Data refreshed daily, snapshot 2026-10-11.