ATRIUMsearch → argument graph
MechanismVideo · 25:03 — 28:00

1X is all-in on pre-training their own models on general internet video data because Neo's human-like form factor allows it to leverage the vast quantity of video data of humans as a training source, which is multiple orders of magnitude larger than any robotics-specific data collection effort.

1X CEO Bernt Børnich explains that the company's decade-long bet is making Neo as human-like as possible so it can be trained on the vast quantity of general video data of humans on the internet, which is orders of magnitude larger than any data collected from teleoperation or sensor-wearing humans. ✦ AI generated

Bernt Børnich · All-In Podcast · 2026-07-29 · original ↗

starts at this moment · 25:03

Elicited by

Take us inside that lab. You are having people in factories wear glasses, wear your hands, and do their tasks over and over again?

The big bet that we made which is this decade long bet in 1x is if you get the robot to be similar enough to a human then you can train on all of the available video data out there of humans. And we're starting to see some very good proof that this is actually working incredibly well. And that's the reason we started the WX World Model Lab because we now finally have the scaling loss on that. And we're seeing that this... If you look at what is needed to actually achieve true intelligence, you need multiple orders of magnitude more data than anyone is even close to collecting over the next few years with egocentric data or with this sensor data. And all of the major breakthroughs that we've seen as far as I'm aware of in AI have been because someone figured out how to use a huge new data source that previously we were not able to use. You unlock some new set of data and now your model capability greatly improves. Our big bet is you have to be able to utilize the general video data out there. And the only way to do that is you have to care about every single tiny detail of the robot to be as close to human as possible.

verbatim transcript · starts at 25:03

Transcript · around this moment

24:47clunky and like so we're increasingly seeing that gathering data with humans just wearing the sensors of the robot in as transparent a manner as possible. So like they should not disturb what you are doing right that's the most useful data to solve kind of baseline dexterity on the robot but even more importantly the big bet that we made which is this decade long bet in 1x is if you get the

25:11robot to be similar enough to a human then you can train on all of the available video data out there of humans. >> Yes. >> And we're starting to see some very good proof that this is actually working incredibly well. And that's the reason we started the WX World Model Lab because we now finally have the scaling loss on that. And we're seeing that this >> take us inside that take us inside the

25:29lab. You are you having people in factories wear glasses, wear your hands, and do their tasks over and over again? Are you working with the micro ones of the world to go do you know real world stuff and outsourcing like unique proprietary data that you can have that other companies don't? How does the world model get built at scale? >> So so so first of all yes we do that and

25:53if you but that's not the main point. So I think ultimately it's very simple right the model is going to be as good as the data. >> Yeah. And if you think about the data pyramid then on the top you have like tele operation data very high quality small fine tuned data set where actually what we do is you will have the operator try to do the task very well and very

26:17fast and they will often fail and then just try again and then we pick the good samples where they did the task as good as a human would right >> you don't need a lot of that data it's just to align your model then you have the data which is what you're talking about with like put the sensors on the human go and gather data. >> Yeah,

26:33>> you have more of that and it's very close to the robot but it's not the robot. The telea is the robot. This is not the robot but it's close. >> Then you have egocentric video from humans point of view. So that is further away from the robot but it's still quite close because the robot hands is the same as human hands and like it looks the same and so it's quite close.

26:54>> And then you have general video data. >> Yes. >> Of the world. or the world in general and of people, right? And because Neo is so similar to a human, we can actually utilize all of that data. Now, the bottom layer in the pyramid, which is this video data, general video data is absolutely ludicrously immense compared to anything else. >> It's YouTube, it's everything. So if you

27:20look at what is needed to actually achieve true intelligence, >> you need multiple orders of magnitude more data than anyone is even close to collecting over the next few years with egocentric data or with this sensor data. >> Got it? >> And all of the major breakthroughs that we've seen as far as I'm aware of in AI have been because someone figured out how to use a huge new data source that

Around this claim