SommBench - Wine Feature Completion (German): leaderboard
Metric: Wine feature completion accuracy (%) with German output: attribute accuracy on SommBench's 1,000 wine profiles with one to three masked attributes (type, country, region, grapes, dryness, body, acidity, sugar, alcohol) completed as structured JSON in the target language, numeric attributes counted correct within 5% error, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-4o | 65 | #333 |
| 2 | Gemini 2.5 Flash | 63 | #237 |
| 3 | GPT-4o Mini | 63 | #588 |
| 4 | Gemini 2.5 Pro | 62 | #145 |
| 5 | GPT-4.1 Mini | 62 | #346 |
| 6 | GPT-4.1 | 62 | #240 |
| 7 | Grok 4 Fast | 61 | #242 |
| 8 | Grok 4 | 59 | #169 |
| 9 | Gemini 2.5 Flash Lite | 59 | #413 |
| 10 | GPT-5 | 56 | #91 |
| 11 | GPT-4.1 Nano | 54 | #716 |
| 12 | GPT-OSS-20B (Medium) | 42 | #499 (GPT-OSS-20B) |
| 13 | GPT-OSS-120B (Low) | 42 | #330 (GPT-OSS-120B) |
| 14 | GPT-OSS-120B (Medium) | 42 | #330 (GPT-OSS-120B) |
| 15 | GPT-OSS-20B (Low) | 36 | #499 (GPT-OSS-20B) |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-feature-completion-german · How It Works · Data refreshed daily, snapshot 2026-10-11.