ATRIUMsearch → argument graph
ClaimVideo · 7:54 — 9:20

No one really cares about prediction of the future; what people actually want is to shape the future through understanding causal mechanisms and counterfactuals — what they need to do now to change the outcome.

Joon argues that prediction alone is insufficient — companies and policymakers need causal models that show them how to intervene and change the future, which is why Simile focuses on randomized control trials and counterfactual data. ✦ AI generated

Joon Sung Park · 20VC · 2026-08-01 · original ↗

starts at this moment · 7:54

Elicited by

How do you think about the chasm of what people say and what people do and how that impacts your models?

My personal hot take here is a lot of observational behavior data and what they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction task. But, my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. No one really cares about what's going to happen in the future unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know, imagine you're a Starbucks, doesn't really help them to know that your Frappuccino sales is going to tank in two quarters. They'll hear that and they'll be like, "What do we do about them? That's terrible." What they want to know is how can we prevent it? What do we need to do now to change the future? And there, what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals.

verbatim transcript · starts at 7:54

Transcript · around this moment

7:40people's values, preferences, and taste. Sort of their subjective half of their brain. That's what we care about. >> I love that. A lot of what people say is different to a lot of what people do. How do you think about the chasm of what people say and what people do and how that impacts your models? >> For sure. So, say to give us real and you know, if

8:04you look at the web data, it is fundamentally data of what people have said, not what they have done. And obviously, things models today are trained uh preliminary mainly on this web data. For us, we actually do collect a lot of behavior data. We collect uh transaction data. We collect observational data. We also partner with our uh customers, uh our vendors to collect some of this data.

8:30But, my personal hot take here is a lot of observational behavior data and what they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction task. But, my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. No one really cares about what's going to happen in the future

8:55unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know, imagine you're a Starbucks, doesn't really help them to know that your Frappuccino sales is going to tank in two quarters. They'll hear that and they'll be like, "What What do we do about them? That's terrible." What they want to know is how can we

9:17prevent it? What do we need to do now to change the future? And there, what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals. So, the kind of data that we care deeply about is a lot of randomized control trials. We actually run a lot of AB testing. We show the models, imagine people have done this versus that. This

9:39is how their behaviors will actually change. That becomes a core part of our training asset. So, this is actually the data collection that goes beyond observational data that Simility collects. >> Is data collection acquisition the hardest element of building simulation models for you? Like if you think about the kind of core pillars for traditional models, it might be compute, algorithms, and data. Is Is data the biggest

Around this claim