Olfactory Perception Benchmark - Smell Identification: leaderboard

Metric: Score (%) on smell identification (pick the odor source, such as a food, of a volatile-compound mixture from four options, 30 questions) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.580
2Claude Opus 4.6 (Max)80
3GPT-5 Pro80
4GPT-5 (High)76.7
5Grok 4.1 Fast73.3
6O4 Mini (High)73.3
7Claude Opus 4.6 (High)73.3
8GPT-5 (Low)73.3
9Grok 3 Mini (Low)73.3
10O3 (High)70
11Claude Opus 4.5 (High)70
12Grok 3 Mini (High)66.7
13GPT-OSS-120B (High)56.7
14Llama 3.3 70B Instruct46.7

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-smell-identification · How It Works · Data refreshed daily, snapshot 2026-10-07.