Olfactory Perception Benchmark: leaderboard

Metric: Overall score (%): unweighted mean of the eight task scores (accuracy for single-answer tasks, multilabel F1 for rate-all-that-apply and receptor activation; the odor similarity of mixtures task, where no model exceeds chance on similar pairs, is included in the mean) over the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)64.4
2Claude Opus 4.6 (High)63.1
3Claude Opus 4.5 (High)62
4GPT-5 Pro61.9
5GPT-5 (High)61.1
6Claude Sonnet 4.559.6
7GPT-5 (Low)59.6
8O3 (High)59.3
9O4 Mini (High)59
10Grok 4.1 Fast58.3
11Grok 3 Mini (High)57.9
12Grok 3 Mini (Low)57.3
13GPT-OSS-120B (High)54
14Llama 3.3 70B Instruct52.7

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark · How It Works · Data refreshed daily, snapshot 2026-10-07.