Olfactory Perception Benchmark - Olfactory Receptor Activation: leaderboard

Metric: Score (%) on olfactory receptor activation (which of 4 to 10 human receptors a molecule activates, 80 questions from M2OR; per-question multilabel F1) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 21 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)51.1
2Claude Opus 4.6 (High)49.6
3Claude Opus 4.5 (High)45.8
4GPT-5 Pro42.5
5O3 (High)42.4
6Grok 3 Mini (High)41.9
7GPT-5 (High)40.8
8O4 Mini (High)40.5
9GPT-5 (Low)40.4
10Claude Sonnet 4.538.4
11Grok 3 Mini (Low)37.2
12GPT-OSS-120B (High)35.6
13Llama 3.3 70B Instruct35
14Grok 4.1 Fast31.1

Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-olfactory-receptor-activation · How It Works · Data refreshed daily, snapshot 2026-10-07.