SommBench - Wine Theory QA: leaderboard

Metric: Wine theory QA accuracy (%), mean over eight prompt languages: accuracy on SommBench's 128 four-option wine theory questions (WSET-style oenology), answered with a single letter, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1GPT-597#91
2Gemini 2.5 Pro96#145
3Grok 496#169
4Grok 4 Fast93#242
5GPT-4o90#333
6GPT-4.190#240
7GPT-OSS-120B (High)85#330 (GPT-OSS-120B)
8GPT-OSS-120B (Medium)84#330 (GPT-OSS-120B)
9Gemini 2.5 Flash Lite83#413
10GPT-4.1 Mini80#346
11GPT-4o Mini80#588
12GPT-OSS-120B (Low)80#330 (GPT-OSS-120B)
13GPT-4.1 Nano73#716
14GPT-OSS-20B (Medium)72#499 (GPT-OSS-20B)
15GPT-OSS-20B (Low)67#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-theory-qa · How It Works · Data refreshed daily, snapshot 2026-10-11.