ATRIUMsearch → argument graph
DataVideo · 21:42 — 23:12

Osmo has digitized 6 billion candidate molecules and 5.43 million human sniffs, making it what Alex believes is the largest olfactory dataset ever built for training AI models.

Alex quantifies Osmo's data scale: billions of enumerated candidate molecules and millions of labeled human sniffs, calling it the largest olfactory AI dataset ever created. ✦ AI generated

Alex Wiltschko · The TWIML AI Podcast · 2026-07-08 · original ↗

starts at this moment · 21:42

Elicited by

What does that mean? Walk us through that process.

We've digitized 5.43 million sniffs. Um, so that's like the largest alactory data set for, you know, the purposes of training AI models, I think, ever. Um, and, uh, that's we had to make all of that from scratch, right? Like there's no scale AI or there's no mechanical Turk for smell.

verbatim transcript · starts at 21:42

Transcript · around this moment

21:22well as safety. Safety is really critical. Um, and uh, that's kind of one of our core data sets. The other thing is we've smelled a lot like so we train people to smell. There's many different protocols for how to like smell something. Hey, is this smell good or bad or intense or not intense? Is this better than the other one? There's like many different ways of doing this and

21:42we've gotten really good in at dialing that in. Uh, and we have people basically internationally that smell. And so we've, uh, I think I want to get the number right. It's probably shifted since I last looked this up like a week ago, but we've digitized 5.43 million sniffs. Um, so that's like the largest alactory data set for, you know, the purposes of training AI models, I think, ever. Um, and, uh, that's we had

22:08to make all of that from scratch, right? Like there's no scale AI or there's no mechanical Turk for smell. like we've had to make that internally, consume it ourselves and generate huge amounts of alactory data for alactory intelligence. >> And that data set ultimately looks like uh a molecule, however you want to represent that, and a set of labels that the human might label that smell. >> It can be broader than that, right? So,

22:32um that's a part of our data set is like we know the molecular structure and then we smell it and label what it smells like. >> But it also might be like a cucumber you buy from the grocery store that we smell, right? Or or we [clears throat] might uh and that's the case of you know analytical annotations it might be a product like a market product.

22:48>> Okay. So a thing and a smelly amount. >> Exactly. And sometimes we don't just smell it. We put it through analytical machinery. Right. So that gets to how you actually where's the real world data come from like you need to use chemical sensors. And so we put huge amounts of data through chemical sensors. And then we also can align that with human sensory labels so that we can begin to

23:08actually relate sensors to human perception which and like that's kind of core to what we do. And when you're in this part of the process where you're manufacturing these molecules um like is there are there known toxicity screens that you can >> Oh yeah. you you have to go through a very rigorous process in Europe, in the US and and worldwide and and there's there's a binder of um of tests you have

23:33to submit and and they're really thorough and they're the right tests, right? So, you know, is this safe on your skin? Is it safe to breathe in? Is this safe for your eyes? Is it safe uh for fish? Because you might flush some of it down the, you know, the toilet or or in the shower drain right after you wash yourself with a with a shampoo. The

Around this claim