ATRIUMsearch → argument graph
DataVideo · 22:08 — 23:38

Osmo built the largest olfactory dataset ever assembled from scratch, since no existing service like Scale AI or Mechanical Turk exists for smell data collection.

Because no crowdsourcing infrastructure exists for scent labeling, Osmo had to build its own pipeline internationally, resulting in what Wiltschko calls the largest olfactory dataset ever created for AI training. ✦ AI generated

Alex Wiltschko · The TWIML AI Podcast · 2026-07-08 · original ↗

starts at this moment · 22:08

Elicited by

What does that mean? Walk us through that process.

we've digitized 5.43 million sniffs. Um, so that's like the largest alactory data set for, you know, the purposes of training AI models, I think, ever. Um, and, uh, that's we had to make all of that from scratch, right? Like there's no scale AI or there's no mechanical Turk for smell.

verbatim transcript · starts at 22:08

Transcript · around this moment

22:08to make all of that from scratch, right? Like there's no scale AI or there's no mechanical Turk for smell. like we've had to make that internally, consume it ourselves and generate huge amounts of alactory data for alactory intelligence. >> And that data set ultimately looks like uh a molecule, however you want to represent that, and a set of labels that the human might label that smell. >> It can be broader than that, right? So,

22:32um that's a part of our data set is like we know the molecular structure and then we smell it and label what it smells like. >> But it also might be like a cucumber you buy from the grocery store that we smell, right? Or or we [clears throat] might uh and that's the case of you know analytical annotations it might be a product like a market product.

22:48>> Okay. So a thing and a smelly amount. >> Exactly. And sometimes we don't just smell it. We put it through analytical machinery. Right. So that gets to how you actually where's the real world data come from like you need to use chemical sensors. And so we put huge amounts of data through chemical sensors. And then we also can align that with human sensory labels so that we can begin to

23:08actually relate sensors to human perception which and like that's kind of core to what we do. And when you're in this part of the process where you're manufacturing these molecules um like is there are there known toxicity screens that you can >> Oh yeah. you you have to go through a very rigorous process in Europe, in the US and and worldwide and and there's there's a binder of um of tests you have

23:33to submit and and they're really thorough and they're the right tests, right? So, you know, is this safe on your skin? Is it safe to breathe in? Is this safe for your eyes? Is it safe uh for fish? Because you might flush some of it down the, you know, the toilet or or in the shower drain right after you wash yourself with a with a shampoo. The

23:48question that I'm curious about is can you derive that from molecular structure or do you have to do it empirically? You can predict it which is super important for how we're so efficient at what we do. Um but then you have to test it physically. It's just the law. Um and it's the right thing to do. Like you just just check right do the experiment and and we do um over and over for all

Related moments