Olfactory Perception Benchmark - Rate-All-That-Apply: leaderboard
Metric: Score (%) on rate-all-that-apply descriptor profiling (select every applicable descriptor from 138 for 100 molecules; per-question multilabel F1) of the 1,010 questions of the Olfactory Perception (OP) benchmark presented with compound-name prompts (the isomeric SMILES variant is reported only in figures), constrained option lists, provider APIs without web search or tools, each model at its stated reasoning effort or thinking budget; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 21 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.5 (High) | 42.2 |
| 2 | Claude Opus 4.6 (High) | 40 |
| 3 | Claude Opus 4.6 (Max) | 38.9 |
| 4 | Grok 3 Mini (High) | 37 |
| 5 | GPT-5 (High) | 36.4 |
| 6 | GPT-5 Pro | 36.1 |
| 7 | Grok 3 Mini (Low) | 36 |
| 8 | Grok 4.1 Fast | 35.5 |
| 9 | Claude Sonnet 4.5 | 34.9 |
| 10 | O3 (High) | 31.5 |
| 11 | GPT-5 (Low) | 30.8 |
| 12 | O4 Mini (High) | 29 |
| 13 | Llama 3.3 70B Instruct | 26.8 |
| 14 | GPT-OSS-120B (High) | 25.1 |
Interactive version: theaggregate.ai/benchmark?slug=olfactory-perception-benchmark-rate-all-that-apply · How It Works · Data refreshed daily, snapshot 2026-10-07.