SommBench - Wine Theory QA (English): leaderboard

Metric: Wine theory QA accuracy (%) with English prompts: accuracy on SommBench's 128 four-option wine theory questions (WSET-style oenology), answered with a single letter, zero-shot, temperature 0, one run; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro98#145
2GPT-598#91
3Grok 498#169
4Grok 4 Fast94#242
5GPT-4o91#333
6Gemini 2.5 Flash91#237
7GPT-4.191#240
8GPT-4.1 Mini90#346
9GPT-OSS-120B (High)88#330 (GPT-OSS-120B)
10GPT-OSS-120B (Medium)86#330 (GPT-OSS-120B)
11Gemini 2.5 Flash Lite85#413
12GPT-4o Mini84#588
13GPT-OSS-120B (Low)81#330 (GPT-OSS-120B)
14GPT-4.1 Nano80#716
15GPT-OSS-20B (Low)75#499 (GPT-OSS-20B)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=sommbench-wine-theory-qa-english · How It Works · Data refreshed daily, snapshot 2026-10-11.