NEWS
Yale AI Maps Scent Mixes in a Space of Words
A Yale-led contest found that word labels predicted scent-mixture similarity better than chemical structure, with a 0.08 error on hidden pairs.
A Yale-led machine learning ensemble predicted how similar two scent blends smell, and the features that worked were words, not raw chemical bonds. Error on a hidden set of 46 pairs fell about 33% to 0.08 on a 0-to-1 scale, according to a paper in the Proceedings of the National Academy of Sciences dated August 4, 2026.
Lead author Vahid Satarifard, a research scientist at Yale’s Human Nature Lab, said English’s thin smell vocabulary had been treated as a weak basis for this job. The winning models used it anyway, and an extra averaging step that kept only odor-word features raised the test-set Pearson correlation to 0.61.
Four Labs Tied, Then Averaged Their Models
The work is the public write-up of the DREAM Olfactory Mixtures Prediction Challenge, organized by Pablo Meyer of IBM Research. Twenty-six teams spent about three months trying to guess how close two mixtures would smell to people, scored on root mean squared error and Pearson correlation.
Organizers merged six datasets from three earlier sniffing studies into 168 single molecules, 731 mixtures, and 507 pair ratings. Distances ran from 0 (no difference) to 1 (as different as the scale allows). Teams could poke a 46-pair leaderboard up to three times a week. Final grading used another 46 pairs that stayed hidden, with 10,000 bootstrap draws. Nineteen teams filed that test set after 159 leaderboard tries.
Four groups tied for first: D2Smell, UMich CASI, ChemSenSim Lab, and belfaction. After the contest they pooled those four with two more models, PL21 and Tamarin, each already under 0.1 RMSE, into a six-model average Satarifard called a machine-learning version of the wisdom of crowds.
THE FOUR TEAMS THAT TIED
- D2Smell: Yale’s Human Nature Lab with partners at KTH, Weizmann, and Cornell Tech, using boosted trees on aroma-pair word labels.
- UMich CASI: University of Michigan statisticians using 138 word labels from the Principal Odor Map, grouped into 12 smell families.
- ChemSenSim Lab: Université Côte d’Azur and Oxford, scoring mixtures with statistical summaries of those same word embeddings.
- belfaction: KU Leuven, feeding molecule-perception embeddings into extremely randomized trees.
On the hidden 46 pairs the average hit a median RMSE of 0.08 and a Pearson correlation of 0.57, beating each of the six parts and beating older baselines that included the Snitz angle, an open Principal Odor Map score, a prior word model, and an aroma-pair model. A later check on 50 newly mixed pairs kept the same ranking. The paper, titled as a semantic-based community model, argues those perceptual distances between odor mixtures can be mapped with high fidelity. Code went public with the models, datasets, and evaluation scripts on GitHub. PNAS also ran the study as a cover paper in mid-August, which Satarifard posted from his account @VahidSatarifard on August 14.
Words Ranked Higher Than Chemical Bonds
SHAP feature ranks, a method that assigns credit to each input, did not crown bond counts or atom types. They crowned compact word labels learned from single molecules, including Principal Odor Map probabilities and aroma-pair tags such as fruity, green, and sulfurous.
For D2Smell, pair-model labels carried the most weight, with a mean absolute SHAP value of 0.11. UMich CASI’s top three were “Fruity,” “Green & Herbal,” and “Nutty, Fatty & Roasted.” ChemSenSim got its biggest lift from absolute-difference summaries of those embeddings, at 0.10. PL21 still leaned on Dragon chemical descriptors, at 0.081, and it was the exception that proves the stack. Three of the strongest models reached half their SHAP importance with fewer than 100 features.
We found that using language features was very powerful in predicting similarity between two scent mixtures. This is interesting because English has a small vocabulary for smell, and we usually describe odors by naming objects. Something ‘smells like flowers’ or like watermelon or like a rotten egg, whereas colors have specific names like ‘green’ or ‘blue.’ That was thought to make semantic odor descriptors a poor basis for predicting how smells relate, but we found the opposite.
Vahid Satarifard, research scientist, Yale Human Nature Lab
When the authors rebuilt the crowd model using only those olfactory word features, Pearson correlation rose 7% to 0.61 on the hidden test and 15% to 0.54 on the 50-pair check. Because the words had been extracted from pure molecules, they wrote, mixture perception may not need a different kind of map from single-molecule smell. The hard step is still turning a structure into a percept. Once that percept is named, comparing two blends looks more like comparing two short descriptions than like overlaying two molecular graphs.
That finding sits next to a quieter one. Several winners did not ignore chemistry. They ingested chemistry that had already been translated into words by earlier models, then did the mixture math in that word space. Dropping the leftover chemical bits helped. The coordinate system that survived is the impoverished lexicon people actually use at the table.
The Color Wheel Smell Never Got
Vision has trichromatic theory. Sound has frequency. Perfume still has “smells like apple pie.” Satarifard put the gap in plain terms: a decade of work can say what one molecule does, but a cup of coffee is dozens or hundreds of molecules, and the field had no shared yardstick for those blends.
THE PATH TO A MIXTURE METRIC
- February 2017: The first DREAM olfaction challenge, reported in Science, asks teams to predict sensory attributes of odor molecules from chemical features and hits eight of 19 word labels, plus intensity and pleasantness.
- August 2023: A graph neural net produces a Principal Odor Map that places single molecules in a shared smell space and, on 400 new odorants, matches a 15-person panel mean better than the median panelist.
- Summer 2024: Team write-ups for the mixture challenge go up on the DREAM Synapse page, shifting the target from one molecule’s adjectives to the distance between two blends.
- December 16, 2025: The consortium posts a preprint of the ensemble and the 50-pair check.
- August 4, 2026: PNAS publishes the peer-reviewed paper, then features it on a later issue cover.
The 2023 map is still chemistry at the front door. Osmo, the company that grew out of the Google Research group behind it, still describes the job as building a map of odor from molecular structure. The DREAM mixture result does not throw that map away. It shows that the useful distances between blends live in the word layer that map emits, which is why a color-wheel metaphor for coffee and pie was always going to be a dictionary first and a periodic table second.
How Accurate Is a 0.08 Error?
On a 0-to-1 scale, 0.08 RMSE means a typical miss is small relative to the full range, and it is about a third below the older scores the authors treated as state of the art. Pearson 0.57, then 0.61 with words only, is a clear lift, not a finished sense of smell. Split-half checks on the human ratings still sit above the models, so the contest did not hit the ceiling of the noses that supplied the labels. Each pair in the new tests also carried 30 human measurements, which is a lot of sniffing for 46 comparisons and still a thin sample by vision-science standards.
THE SCOREBOARD ON HIDDEN PAIRS
| What was scored | Figure | Set |
|---|---|---|
| Training mixture pairs | 507 | Three prior sniffing studies |
| Hidden test pairs | 46 | DREAM holdout |
| New validation pairs | 50 | Built after the contest |
| Ensemble RMSE | 0.08 | Hidden test |
| Ensemble Pearson r | 0.57 | Hidden test |
| Word-only Pearson r | 0.61 | Hidden test |
| Word-only Pearson r | 0.54 | 50-pair validation |
The hidden set was not a random grab bag. It mixed three kinds of pairs: random blends, hand-built blends meant to smell close, and blends matched by the Snitz angle. Models did better on the hand-built and Snitz pairs than on the random ones, which is a polite way of saying they learned the designers’ structure more easily than chance collisions. That is useful in a lab and less comforting in a kitchen, where nobody designed the steam.
Parkinson’s Scent Is a Mixture Problem
Satarifard’s pitch for a use case is disease, not dessert. Parkinson’s and some cancers have been reported to carry odor signatures, he said, and one long goal is to treat smell as a biomarker, with gear years from now that watches a person’s odor and flags a change that needs a screen. Nicholas Christakis, Sterling Professor of Sociology and Natural Science at Yale and director of the Human Nature Lab, added a non-clinic use: body scent is itself a complex mixture, and he suspects it shapes how people deal with one another. The lab had already posted a social-chemosignaling fellowship along those lines in late 2025.
Body scent is also a complex mixture of odors, and we suspect that it plays an important role in human social interactions.
Nicholas Christakis, director, Yale Human Nature Lab
The Parkinson’s claim is older than this model. In 2019, chemists working from sebum on the upper back, and from a “Super Smeller” named Joy Milne, reported volatile biomarkers of Parkinson’s from sebum in 43 patients and 21 controls, then checked an extra 31 people, pointing to shifts in perillic aldehyde and eicosane. That is a chemical signature, not a word-space distance. A mixture metric helps only if a later system can turn a patient’s volatiles into the same kind of labeled blend the DREAM models know how to compare. Nothing in the PNAS paper does that translation, and Satarifard did not claim it does. He called it an ultimate goal, several years out.
A 10-Molecule Mix Is Not Coffee
Christakis, posting as @NAChristakis on August 6, 2026, went further than the journal text and said the work lays a foundation for digital olfaction, machine-readable smell metrics, next-generation electronic noses, and odor teleportation.
Olfaction lacks a quantitative framework similar to vision and hearing.
In new #HNL work in @PNASNews, we show that odor mixture perception can be predicted. This lays a foundation for digital olfaction, machine-readable smell metrics, next-generation electronic noses, and…
— Nicholas A. Christakis (@NAChristakis) August 6, 2026
The paper’s own limits sit next to that list. Test and validation pairs used molecules from a closed palette, and the authors say it is still open whether the models hold for chemicals they have never seen. The new human tests used 92 mixtures that formed 46 pairs, each mix built from 10 non-overlapping components drawn from an intensity-matched set. People judged them in triangle tests, picking the odd vial among three. That is a clean psychophysics design. It is not a bakery.
WHAT THE MAP STILL CANNOT DO
- New chemicals: Generalization past the 168-molecule training list is untested in the published sets.
- Real food: Coffee and pie were the public examples; the graded items were 10-component lab blends.
- Human ceiling: Split-half reliability on the sniffers still leaves room the ensemble did not close.
- A device: No sensor, no phone accessory, and no protocol for sending a smell over a network appears in the paper.
The public conversation around the cover was thin, which fits a methods result more than a product launch. The authors are already talking about teleporting odor. The numbered evidence is still a 0.08 error on 46 hidden pairs of 10-molecule mixes, scored in a space of words English barely has names for, and that is the part that should travel with the headline.
Frequently Asked Questions
What is a triangle test in a smell-mixture study?
Panelists get three vials, two holding one mixture and one holding the other, and must pick the odd smell. In this challenge’s new tests, all three vials were intensity-matched so people could not win by noticing loudness, and the two mixtures in a pair shared no components.
What did the 2017 DREAM olfaction challenge actually predict?
Twenty-two teams received 4,884 physicochemical features on 476 molecules that 49 people had sniffed, and the best models predicted intensity, pleasantness, and eight of 19 word labels, including garlic, fish, sweet, fruit, burnt, spices, flower, and sour.
What is the Snitz angle the new models had to beat?
It is an older mixture-distance score that treats each blend as a vector of 21 Dragon chemical descriptors and measures the angle between two of those vectors. Organizers also built 20 of the hidden-set mixtures to be close in that angle, then watched whether newer models could beat the metric on its own turf.
How were the 138 Principal Odor Map labels used here?
UMich CASI and others took 138 word-probability scores the 2023 map emits for each molecule, averaged or summarized them inside a mixture, then compared those summaries across a pair. Those probabilities, not the raw graph of atoms, were the features SHAP ranked at the top.
Does this paper mean a phone can send a smell?
No. The published system scores distances between known lab mixtures and does not encode, compress, or replay an odor. Christakis named odor teleportation as a long-term foundation, while the paper’s concrete next experiments are in-silico olfactory metamers, blends that should smell alike on the new metric and can then be mixed and sniffed.
The GitHub release already lets other labs rerun the six models on the 50 new pairs. Until someone scores a real kitchen blend against that yardstick, the map’s unit of distance is still a word-labeled 10-molecule vial, not the steam off a pie.
Disclaimer: This article is news reporting on a published machine learning study and related chemical research, and it is for information only. It is not medical advice, a diagnostic method, or a recommendation to use smell, sebum tests, or any electronic nose to screen for Parkinson’s disease, cancer, or any other condition. Anyone who is worried about a health change should talk with a licensed physician before acting on odor-based claims. Figures and study status come from the PNAS paper, the DREAM materials, and the cited biomarker study as those sources stood on their publication dates, and later work may change the numbers and the limits.
-
NEWS2 months agoMicrosoft’s 95.95% Cyber AI Score Skips Its Own Leaderboard
-
NEWS3 weeks agoG20 Cheers AI Investment After Bailey’s Frontier Cyber Warning
-
NEWS3 months agoGoogle Pixel 10 at Rs 64,649 on Amazon India: Buy Now or Wait?
-
SPORTS3 months agoFree Live Sports Streaming in 2026: What to Watch Without Cable
-
NEWS3 weeks agoThe Yankees Bet Aaron Judge Can Skip the Minors
-
NEWS4 months agoMusk to Dell: How AI Split the World’s 10 Richest in 2026
-
NEWS5 months agoSamsung Galaxy Tab A11 Plus Lands as 2026’s Top Budget Pick
-
BUSINESS2 years agoRetail Banking vs Commercial Banking: A Comprehensive Guide
