Olfactory Perception Benchmark - Odor Intensity: leaderboard

Metric: Score (%) on odor intensity paired comparison (which of two molecules smells more intense, 175 pairs from Keller et al. 2017 ratings; any-overlap accuracy) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)74.9
2Claude Opus 4.6 (High)71.4
3GPT-5 Pro71.4
4Claude Opus 4.5 (High)71.4
5GPT-5 (Low)70.9
6O4 Mini (High)69.1
7Llama 3.3 70B Instruct68
8O3 (High)68
9Grok 3 Mini (High)68
10Claude Sonnet 4.566.9
11Grok 4.1 Fast66.9
12GPT-5 (High)66.3
13Grok 3 Mini (Low)66.3
14GPT-OSS-120B (High)65.1

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-odor-intensity · How It Works · Data refreshed daily, snapshot 2026-10-07.