Papers
Topics
Authors
Recent
Search
2000 character limit reached

Benchmark for Assessing Olfactory Perception of Large Language Models

Published 8 Mar 2026 in cs.CL and cs.AI | (2604.00002v1)

Abstract: Here we introduce the Olfactory Perception (OP) benchmark, designed to assess the capability of LLMs to reason about smell. The benchmark contains 1,010 questions across eight task categories spanning odor classification, odor primary descriptor identification, intensity and pleasantness judgments, multi-descriptor prediction, mixture similarity, olfactory receptor activation, and smell identification from real-world odor sources. Each question is presented in two prompt formats, compound names and isomeric SMILES, to evaluate the effect of molecular representations. Evaluating 21 model configurations across major model families, we find that compound-name prompts consistently outperform isomeric SMILES, with gains ranging from +2.4 to +18.9 percentage points (mean approx +7 points), suggesting current LLMs access olfactory knowledge primarily through lexical associations rather than structural molecular reasoning. The best-performing model reaches 64.4\% overall accuracy, which highlights both emerging capabilities and substantial remaining gaps in olfactory reasoning. We further evaluate a subset of the OP across 21 languages and find that aggregating predictions across languages improves olfactory prediction, with AUROC = 0.86 for the best performing language ensemble model. LLMs should be able to handle olfactory and not just visual or aural information.

Summary

  • The paper introduces the first ground-truth benchmark for LLM olfactory knowledge, evaluating 21 model configurations on 1,010 questions spanning odor classification, descriptors, ratings, mixtures, receptors, and identification.
  • The paper finds Claude Opus 4.6 performs best at 64.4% overall, while compound-name prompts outperform isomeric SMILES by 2.4–18.9 percentage points, indicating that models rely more on lexical associations than molecular structure.
  • The paper shows mixture-level odor similarity remains largely unsolved, but multilingual and cross-model ensembles improve descriptor prediction to AUROC 0.896, suggesting a practical path for stronger olfactory systems.

Motivation and positioning

Olfaction presents an unusual challenge for computational modeling: unlike vision or audition, where physical stimuli map relatively predictably onto perceptual dimensions, odor perception arises from a poorly understood interplay between molecular structure and receptor biology. Predicting how a molecule smells from structure alone has resisted general solution, even as specialized graph neural networks have approached human-level performance on descriptor prediction (2604.00002). Prior evaluations of LLMs on sensory alignment explicitly excluded olfaction, and existing olfaction-focused work either measured embedding-space alignment through subjective descriptions (SNIFF AI, 27.5% success) or semantic similarity of odor words rather than factual accuracy. The Olfactory Perception (OP) benchmark fills this gap: it is, to the authors' knowledge, the first structured question-answering evaluation of whether general-purpose LLMs possess correct factual knowledge about olfactory properties, with ground-truth answers drawn from peer-reviewed datasets.

Benchmark design

The benchmark comprises 1,010 multiple-choice questions across eight categories, each grounded in established olfactory science resources:

Category Task type n Source
Odor Classification (OC) Binary odorous/odorless 175 Mayhew et al.
Odor Primary Descriptor (OPD) 4-way choice, 29 descriptors 175 IFRA glossary
Odor Intensity (OIn) / Pleasantness (OPl) Paired comparison + 0–100 rating 175 each DREAM Challenge I
Rate-All-That-Apply (RATA) Multilabel from 138 descriptors 100 GS-LF integrated dataset
Odor Similarity of Mixtures (OS) 4-bin ordinal similarity 100 Snitz, Bushdid, Ravia datasets
Olfactory Receptor Activation (ORA) Multilabel over human OR gene IDs 80 M2OR
Smell Identification Test (SIT) 4-way source identification 30 Leibniz-LSB@TUM database

A central methodological feature is dual prompting: every question is posed twice, once using isomeric SMILES and once using common compound names. Because odor perception is stereospecific—(R)-carvone smells of spearmint while (S)-carvone smells of caraway—the contrast between representations probes whether models reason structurally about molecules or retrieve lexical associations. Scoring uses any-overlap accuracy for single-answer tasks and per-question multilabel F1 for RATA and ORA, with Pearson correlations computed for the continuous-rating tasks.

Overall results

Twenty-one model configurations spanning six providers were evaluated without retrieval or tool use. Claude Opus 4.6 (max) achieves the best overall score at 64.4%, followed by Claude Opus 4.6 (high) at 63.1%, GPT-5.2 Pro at 62.2%, and Claude Opus 4.5 at 62.0%. All scores substantially exceed per-task chance baselines but remain far from ceiling. A capability hierarchy is evident: Anthropic and OpenAI frontier models occupy the top positions; Gemini 2.5 Pro (58.7–59.7%) and Grok (57.3–58.3%) form a mid-tier; DeepSeek Reasoner reaches 56.6–58.5%; and Llama 3.3 70B trails all proprietary systems at 52.7%.

Performance stratifies sharply by task difficulty. Simple tasks yield strong results—up to 92.0% on OC and 80.0% on OPD—whereas intermediate tasks cap at 42.2% F1 (RATA) and 35.0% accuracy (OS). Among hard tasks, SIT reaches 80.0% for three models, aided by food-related world knowledge, while ORA peaks at only 52.8%. Question-level analysis shows that 47.0% of single-label questions are solved by every model while 13.6% are solved by none; OS alone accounts for 43 universally unsolved questions.

Representation format dominates performance

The most consequential finding concerns the dual-prompting manipulation. Compound-name prompts outperform isomeric SMILES for all 21 configurations, with gains from +2.4 to +18.9 percentage points (mean ≈ +7). Llama 3.3 70B shows the largest disparity (33.8% → 52.7%), barely exceeding chance under SMILES, while frontier reasoning models retain over 92–95% of their name-based scores under SMILES. The gap also varies by task: it is smallest for OC (odorousness correlates with inferable properties such as volatility and molecular weight) and largest for OIn, OPl, and SIT, where identifying the compound is decisive. The authors interpret this pattern as evidence that current LLMs access olfactory knowledge primarily through lexical association rather than structural molecular reasoning—a claim that, if sustained across future models, implies that apparent olfactory competence may not transfer to novel or unnamed compounds.

Reasoning budget yields diminishing returns

Systematic variation of reasoning budgets produces consistent but modest gains: no configuration improves by more than roughly two percentage points with extended deliberation (e.g., GPT-5 low→high: 59.6%→61.1%; Gemini 8K→32K: 58.7%→59.7%). DeepSeek Reasoner shows a non-monotonic pattern, suggesting task-dependent optima. This contrasts with chemistry benchmarks such as ChemIQ, where extended reasoning yields substantial gains, and the authors attribute the difference to the more constrained nature of olfactory knowledge relative to general chemical reasoning.

Fine-grained failure modes

Continuous-rating correlations reveal graded competence: best models reach r≈0.55r \approx 0.55 for intensity (approaching specialized-model performance), r≈0.60r \approx 0.60 for pleasantness, but only r≈0.35r \approx 0.35 for mixture similarity, confirming that integrating perceptual information across molecules is the weakest capability.

Two qualitatively distinct failure mechanisms emerge. For OS, all models rely on molecular overlap as a proxy for perceptual similarity: accuracy reaches ~85% when similar mixtures share many molecules but falls near 0% when they share few—even though 36% of ground-truth "Strongly Similar" pairs share zero molecules. Claude models assign "Slightly Dissimilar" to 72–96% of mixtures regardless of ground truth, and no model exceeds chance on combined Similar categories. The authors conclude that mixture-level olfactory similarity is fundamentally beyond current LLM capabilities.

For ORA, difficulty is bimodal: models either possess receptor–ligand knowledge or lack it entirely. A striking case study concerns the hOR2W1_D296N variant: Claude Opus 4.6 never predicts this receptor across all 24 ground-truth appearances, consistently invoking an incorrect loss-of-function hypothesis about the mutation, whereas GPT-5.2 Pro correctly identifies it in 19 of 24 cases—despite starting from the same initial misconception, which it self-corrects during chain-of-thought. This single knowledge gap largely explains GPT-5.2 Pro's ORA advantage. Per-label RATA analysis shows descriptors tied to identifiable functional groups (sulfurous, floral, fruity) achieve high F1, while holistic descriptors such as spicy, fresh, and tropical are nearly unsolvable; in examined reasoning traces for "spicy," the term never appears as a candidate.

An additional finding concerns safety alignment: Claude Opus 4.6 refuses odor-classification questions about hazardous compounds such as Tabun, costing it correct answers, while other model families answer them all correctly. The authors note that nerve-agent odor detection constitutes legitimate toxicology knowledge relevant to protective equipment design, illustrating a tension between safety filtering and scientific evaluation.

Multilingual evaluation

Translating the RATA task into 21 languages across six language families reveals that English achieves the highest mean F1, followed closely by French, Spanish, and Russian, with non-Indo-European languages (Korean, Chinese, Swahili) clustering lower. More notably, aggregating predictions across languages improves performance: a cross-language vote ensemble reaches AUROC = 0.828 for DeepSeek Reasoner (32K), the best single-model multilingual ensemble (Gemini 2.5 Pro) reaches 0.864, and a cross-model ensemble pooling seven models across all languages attains AUROC = 0.896. This indicates that different languages provide complementary olfactory knowledge—an observation consistent with cross-linguistic variation in odor vocabularies—and suggests ensemble strategies as a practical route to improved olfactory prediction.

Limitations

The authors are explicit about several constraints. Ground-truth labels reflect agreement with established datasets rather than mechanistic understanding or experimentally verified predictions. Discrete response formats prioritize reproducibility but under-weight nuanced free-form descriptions; results remain sensitive to prompting and output formatting despite fixed templates. Multilingual variants were produced via automated translation with validation, which may introduce subtle shifts in connotation. Coverage omits concentration effects, temporal dynamics, contextual modulation, and individual/cultural variation in perception. Finally, the benchmark's reliance on curated descriptor vocabularies means saturated items (nearly half of single-label questions solved by all models) may mask brittleness rather than genuine competence.

Conclusion

The OP benchmark provides the first systematic, ground-truth-based evaluation of olfactory factual knowledge in LLMs, establishing that frontier models encode meaningful but incomplete olfactory information (best overall accuracy 64.4%), that this knowledge is accessed predominantly through lexical rather than structural pathways, and that mixture-level perceptual integration remains essentially unsolved. The dual-prompting methodology offers a reusable instrument for distinguishing structural reasoning from name memorization, and the multilingual results open a concrete question the paper leaves unanswered: whether training or inference-time language aggregation can be exploited deliberately to improve olfactory prediction, and whether hybrid LLM–cheminformatics architectures can close the representation gap that currently separates compound names from molecular structure.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 3 tweets with 23 likes about this paper.