ATRIUMsearch → argument graph
Video · 2026-08-01 · 1h 5m · 6 moments

The Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO

✦ AI generated

timeline · colored by role

01
Example

LLMs trained on human behavioral data can be probed to extract realistic human behaviors, enabling simulation of entire lived experiences.

Jun Park explains how LLMs trained on human behavioral data can be probed to extract realistic human behaviors, leading to the 'Smallville' simulation where 25 NPCs autonomously woke up, did routines, formed relationships, and self-organized a Valentine's Day party.

transcript

Jun Park: So, this was 2023. we had this idea that large language models are often used for simpler tasks like classification, simple generation but we thought that these models actually had a lot more potential. One of the early observations that we made was that these models are trained on so much of human behavioral data, sentiment data that were expressed on the web. So if you poke at them sort of at the right angle, you could actually extract a lot of realistic human behaviors out of them. I thought that was really interesting and it was also particularly interesting in that it was domain agnostic. So if you look at the literature in computer science for many decades we've always had the vision of creating agents that are meant to be generalizable that are meant to really be able to act like human in any environment. And my mind went to well maybe we have that opportunity here. So what we ended up doing was well if we were to fast forward many years into doing this what would be the most ambitious vision that we might have and that was creating entire lived experience of a town. So the idea here was we would make a game town and we would populate it with 25 NPCs. So non-playable characters except these characters would actually wake up in the morning, do their routines, go to work, have relationships and do all that. They would actually remember their interactions. They would actually plan their days. And some of the surprising things you end up seeing was the simulation itself was set the day before Valentine's Day and you actually see these agents come together, have parties like self-organized. So they would actually plan parties, they would decorate the cafe and so forth.

rebuts · 2

02
Claim

Simile creates a foundation model of human behavior — not of rationality — that is designed to make the same mistakes and biases humans do, representing the subjective half of the brain.

Jun Park distinguishes Simile from frontier model companies: Simile's models are designed to replicate human irrationality, bias, and subjectivity — not superhuman intelligence — to accurately simulate how real people make decisions and mistakes.

transcript

Jun Park: So the way we see it is if you look at large language model companies today, fundamentally the task they have at hand is to create super rational intelligent machines that are good at coding, that are good at natural sciences and mathematics. Simile doesn't really care about any of those. What we care about is if we have a person make a mistake in this context, we want our models to make the same kind of mistake. We want our models to be biased in the same way humans are. In a way, we want to be a representation of people's values, preferences, and taste, sort of their subjective half of their brain. That's what we care about.

explains mechanism · 2

03
Claim

The real market need is not prediction of the future but causal counterfactuals — understanding how to shape the future.

Jun Park argues that while observational data is good for prediction, nobody actually cares about prediction alone — what enterprises and decision-makers actually want is causal understanding of counterfactuals: if we do X instead of Y, what happens? This requires randomized control trials and AB testing data, not just web data of what people said.

transcript

Jun Park: But my personal hot take here is a lot of observational behavior data. What they're amazing at is actually helping you create a correlation of the observation and what could happen in the future. Good for prediction task. But my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. No one really cares about what's going to happen in the future unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know imagine you're a Starbucks doesn't really help them to know that your fraction of sales is going to tank in two quarters. They'll hear that and they'll be like what do we do about them? That's terrible. What they want to know is how can we prevent it? What do we need to do now to change the future? And there what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals. So the kind of data that we care deeply about is a lot of randomized control trials. We actually run a lot of AB testing. We show the models. Imagine people have done this versus that. This is how their behaviors would actually change. That becomes a core part of our training asset.

provides context · 1

04
Definition

AI companies of this generation need an interesting, defensible data strategy — and the key is sourcing representative everyday people and asking the right questions to capture their fundamental nature.

Jun Park outlines Simile's data strategy: they don't go after expert programmers or scientists but everyday people, ensuring demographic representativeness. The critical data challenge is two-fold: sourcing the right people and asking the right experiments that get at the core of who they are, including their life stories and hardest decisions.

transcript

Jun Park: My fundamental thesis here is for AI companies of this generation you need to have an interesting data strategy that's going to be defensible and for us really the data collection challenge comes from two angles one is actually sourcing people sourcing people here is a little bit different than what other language model companies might consider to be their people or their population. We don't go after these expert programmers or expert scientists. We go after people like us like everyday people living their everyday life. That's what we care about is are they representative? Do we actually have the same representation of people as we do in the world that we live in? And then actually asking the right questions to these people. What are the experiments? What are the questions that actually get at the fundamental core nature of who they are? Some of the questions we actually ask at the start of our data collection at times is actually saying something like tell us the story of your life. Where did you grow up? What did you experience? What were some of the hardest problems that you had to tackle or decisions you had to make? Tell us a lot about these people.

05
Mechanism

Simulation's data flywheel is even stronger than coding agents' because the world itself is the ground truth, generating millions of daily hypotheses that can be validated against real events.

Jun Park explains that unlike coding agents which have a clear reward signal (accept/reject), simulation has a more powerful mechanism: every day the world provides ground truth, allowing Simile to generate tens of thousands of hypotheses, map them to date-specific predictions, and check which percentage came true — creating a compounding learning loop.

transcript

Jun Park: It might be easy to look at simulation as a field and say well where are you going to get the reward? Because fundamentally all the things that you're trying to predict is happening in the future. It's going to be hard to validate. It is true. At the same time I actually think simulation has even better mechanism which is the world is our ground truth. We live in the ground truth world. So what we can do is every single day we can be generating tens of thousands of hypothesis. Each hypothesis is mapped onto an end statement. If this happens, we know whether we can validate the simulation to be right or wrong. And we're basically watching the world every day seeing which of those hypotheses are answerable at what time. And we can basically say a month goes by, we generated a million hypothesis, x percentage of them came true. This is the best way to learn about the world.

extends · 4

06
Prediction

Simulation is the 'GPU of intelligence' — not a single super-intelligent CPU-like model, but many diverse, human-flawed models whose collective emergent phenomena produce the most valuable insights.

Jun Park draws an analogy: today's frontier models are like a CPU of intelligence — one very smart, rational unit. Simulation, by contrast, is the GPU of intelligence — many individually imperfect models that, when combined, produce emergent collective phenomena that can model society, policy outcomes, and complex market dynamics.

transcript

Jun Park: What I see today that's prominent in AI space is what I consider to be the CPU of intelligence unit. You have this one language model that's really large that's very smart that can do very complex reasoning tasks that's like CPU. What I see coming and what I think simulation as a field can offer is the GPU of intelligence unit. As I mentioned before, Simile does not care about creating really smart super intelligent machines. What we care about is creating models that are as smart as we — I fail at a lot of things. I want to make sure that the model that represents me fails the same way. But the beautiful part about people is individually we have so much diversity, so much different takes in our world that makes individuals so interesting. But also when they come together as a large collective, the emerging phenomena that we're able to draw out is some of the most wonderful thing that we can see in our world. Creating a society, creating an amazing process that actually allows us to make all these achievement — can we actually replicate that in simulation?

explains mechanism · 3supports · 2

Highlight slides
Related episodes