Olfactory Perception Benchmark - Odor Pleasantness: leaderboard

Metric: Score (%) on odor pleasantness paired comparison (which of two molecules smells more pleasant, 175 pairs; any-overlap accuracy) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (High)74.9
2Claude Opus 4.6 (Max)74.3
3O4 Mini (High)73.7
4Grok 3 Mini (High)73.7
5Claude Opus 4.5 (High)73.1
6Grok 3 Mini (Low)72.6
7Llama 3.3 70B Instruct72
8GPT-OSS-120B (High)72
9Claude Sonnet 4.571.4
10GPT-5 (High)71.4
11GPT-5 (Low)71.4
12GPT-5 Pro70.9
13Grok 4.1 Fast70.3
14O3 (High)70.3

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-odor-pleasantness · How It Works · Data refreshed daily, snapshot 2026-10-07.