The real market need is not prediction of the future but causal counterfactuals — understanding how to shape the future.
Jun Park argues that while observational data is good for prediction, nobody actually cares about prediction alone — what enterprises and decision-makers actually want is causal understanding of counterfactuals: if we do X instead of Y, what happens? This requires randomized control trials and AB testing data, not just web data of what people said. ✦ AI generated
Jun Park · 20VC · 2026-08-01 · original ↗
starts at this moment · 8:04
“A lot of what people say is different to a lot of what people do. How do you think about the chasm of what people say and what people do and how that impacts your models?”
But my personal hot take here is a lot of observational behavior data. What they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction task. But my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. No one really cares about what's going to happen in the future unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know imagine you're a Starbucks doesn't really help them to know that your fraction of sales is going to tank in two quarters. They'll hear that and they'll be like what do we do about them? That's terrible. What they want to know is how can we prevent it? What do we need to do now to change the future? And there what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals. So the kind of data that we care deeply about is a lot of randomized control trials. We actually run a lot of AB testing. We show the models. Imagine people have done this versus that. This is how their behaviors would actually change. That becomes a core part of our training asset.
verbatim transcript · starts at 8:04
8:04at the web data it is fundamentally data of what people have said not what they have done and obviously large language models today are trained prelim uh preliminary um mainly on this web data. For us we actually do collect a lot of behavior data. We collect uh transaction data. We collect observational data. We also partner with our um customers uh our vendors to collect some of this
8:29data. But my personal hot take here is a lot of observational behavior data. What they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction task. But my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. no one really cares about what's going to happen in the future
8:55unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know imagine you're a Starbucks doesn't really help them to know that your fraction of sales is going to tank in two quarters. They'll hear that and they'll be like what do we do about them? That's terrible. What they want to know is how can we prevent
9:17it? What do we need to do now to change the future? And there what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals. So the kind of data that we care deeply about is a lot of randomized control trials. We actually run a lot of AB testing. We show the models. Imagine people have done this versus that. This
9:39is how their behaviors would actually change. That becomes a core part of our training asset. So this is actually the data collection that goes beyond observational data that similarly collects. Is data collection acquisition. The hardest element of building simulation models for you like if we think about the kind of core pillars for traditional models it might be compute algorithms and data is is data the biggest challenge for you data
10:02is an important piece of simil for sure u my fundamental thesis here is for AI companies of this generation you need to have an interesting data strategy that's going to be defensible and for us really the data collection challenge comes from two angles one is actually sourcing people sourcing people here is a little bit different than what other language model companies might consider to be their people or their population. We
- ·Observational data excels at correlation and prediction
- ·But no one cares about prediction alone
- ·Enterprises want to shape the future, not forecast it
- ·Starbucks doesn't need to know sales will tank
- ·They need to know how to prevent it
- ·Causal mechanisms answer: if we do X vs Y, what happens?
- ·Randomized control trials and AB tests are the core asset
- ·Show models what happens when people do this vs that
- ·Counterfactual behavior data, not just stated preferences