Olfactory Perception Benchmark - Odor Classification: leaderboard

Metric: Score (%) on odor classification (odorous or odorless, 175 molecules, balanced; any-overlap accuracy) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)92
2GPT-5 Pro92
3Claude Opus 4.5 (High)92
4Claude Opus 4.6 (High)91.4
5GPT-5 (Low)90.3
6GPT-5 (High)89.7
7Claude Sonnet 4.589.1
8O3 (High)89.1
9Grok 4.1 Fast88.6
10O4 Mini (High)88.6
11Llama 3.3 70B Instruct83.4
12GPT-OSS-120B (High)82.9
13Grok 3 Mini (High)81.7
14Grok 3 Mini (Low)81.7

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-odor-classification · How It Works · Data refreshed daily, snapshot 2026-10-07.