ATRIUMsearch → argument graph
Video · 2026-07-28 · 42m · 6 moments

Solving the Hardest Problem in Robotics | World Labs with a16z

✦ AI generated

timeline · colored by role

01
Claim

Spatial intelligence—the ability to generate, understand, reason with, and interact with physical or virtual spaces—is the next frontier of AI, and building large world models is the means to achieve it.

Fei-Fei Li introduces World Labs as a frontier model lab building spatial intelligence through large world models.

transcript

Fei-Fei Li: World Lab has is a two-year-old startup. I think we should just recognize it's a frontier model lab. We are building the next frontier of AI which is what we call spatial intelligence. And spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces whether it's physical or virtual. And of course a means to an end towards spatial intelligence is building large world models.

extends · 1

02
Mechanism

The unique bottleneck in general-purpose robotics is the lack of data for training and evaluation, which can be solved by a 'real-to-sim-to-real' pipeline that aligns digital worlds with physical environments so that simulation data replaces real-world data at scale.

Yunzhu Li describes the data bottleneck in robotics and explains Scenix's approach of mapping real environments into aligned digital worlds for scalable training and evaluation.

transcript

Yunzhu Li: Throughout my career, my goal has been very simple: trying to help the robots better perceive and interact with the physical world. So I'm a very practical person. I want my robot to work in the real physical environments. So for Scenix the unique opportunity we see is that there has been a lot of bottlenecks right now we see faced by the developments of general purpose robots especially around training and also around evaluations. So as we are developing what we call a real to sim to real pipeline. We want to map the real environments into the digital world that has the best alignments with the real environments. By alignments we mean that whatever happens in the digital world is also going to happen in the real environments such that we can replace all the data, all the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world.

explains mechanism · 1extends · 2supports · 1

03
Prediction

A foundation model for robotics will be an omni-model that takes multimodal input—including actions as input (a forward simulator predicting environment changes) or actions as output (a policy model)—and serves as a backbone fine-tuned for specific robotic applications.

Fei-Fei Li and Yunzhu Li describe the foundation model for robotics as an omni-model that processes actions as inputs or outputs, acting as both a simulator and a policy backbone.

transcript

Fei-Fei Li and Yunzhu Li: World Lab is building a foundation model. As you know, we're building a base model and as the technology has been evolving, some of the most exciting base models are omni-models, right? They take multimodal input, they have multimodal outputs. And what is a foundation model for robotics? It's very likely going to involve actions. It's very likely going to involve the output of actions in addition to the state of the world and we're definitely not ruling this out. [...] For the foundation models it's essentially needs to be a multimodal model. So it has to take into account frame, text, image, depth and different kind of modalities and action is a very important parts of that modalities. So if you think about using actions as inputs that is essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action. When the action is output this is essentially a policy model that is trying to predict given a specific goal what should be the action you take in the real environment to get you closer to that goal. So this kind of omni models actually can benefit a lot and provide huge amount of values for the robotics communities and this can also acts as a backbone for you to fine-tune into specific robotic applications.

explains mechanism · 1extends · 1

04
Mechanism

The key role simulation plays that real-world data cannot is counterfactual reasoning—playing out events that haven't happened or cannot happen—which is essential because robotics cannot possibly collect enough real-world data for all scenarios.

Fei-Fei Li argues simulation is not a binary alternative to real data but serves a unique role: counterfactual reasoning. She cites Waymo's billions of hours of simulation as a real-world example.

transcript

Fei-Fei Li: There isn't a binary choice between simulation or no simulation. All this come together to make robotics work. Think about human intelligence. We do a lot of simulation in our head. Why? There's a very important role simulation plays that real world data doesn't play which is counterfactual reasoning—you play out events that hasn't happened or cannot happen or you don't have enough data to make it happen in real world. And while you play it out, you learn how to act in it. Humans do this all the time. [...] Here's a real life example: the industry of self-driving cars. Waymo has officially said they use billions of hours of simulation and actually Waymo is more simulation-heavy than just real world data heavy. So these are real examples and as you know cars are the simplest kind of robots. So clearly simulation plays a huge role in robotic learning.

explains mechanism · 3supports · 1

05
Mechanism

Simulation provides two distinct benefits for robotics: reliability through systematic randomization covering the full state space, and efficiency through controllable speed-up of robot behaviors that is impossible in the physical world.

Yunzhu Li specifies simulation's two benefits: reliability via systematic randomization of all variables a robot might encounter, and efficiency by training robots at accelerated speeds that physical constraints prevent.

transcript

Yunzhu Li: Simulation can provide two levels of benefits. The first one is reliability and the second one is efficiency. For reliability, if you're thinking about a robotic system working reliable in the real environments, you need data to provide systematic coverage of all the state space and the variations that robots might encounter. That's how you can learn to be robust. So with simulation, you can do systematic randomizations and control the variations of lighting, frictions, geometries, object types and all different kind of physical parameters to make sure you have sufficient coverage of the state space. So this is what can give the robotic systems reliability. And second is about efficiency. Right now many people are doing teleoperation. You're collecting the data at a speed that is actually slower than human actually doing the task. But for many of our clients human speed is not good enough. They want faster than human speeds. For the robot to move faster, it's not as simple as just drive the robot faster because the gravity doesn't change. But in simulation, you can do systematic speed up of the robots behaviors to train the robots such that it considers all the dynamics changes of the environments.

explains mechanism · 2supports · 1

06
Prediction

Human-level power efficiency in robotics will take a very long time to achieve because a working robot is always a system problem integrating hardware, software, and countless details, and we should take a measured, realistic approach rather than over-promising.

Yunzhu Li argues that human-level efficiency in robotics is far off because every working robot is a complex system, though progress is accelerating. Fei-Fei Li adds that even LLMs cannot match the 30-watt efficiency of the human brain.

transcript

Yunzhu Li and Fei-Fei Li: I think it's going to take a very long time. If you really think about robots in the real environment, in the end it will always be a system. Every working robot in the real environment is a system. You need to be very mindful and thoughtful about how the systems are coming together: the hardware, the software, the brain, even down to the details of what's the friction coefficients of your fingers. So there's a lot of things you have to consider to make these things a reality and it will take iterations. [...] We also have to be calibrated about our predictions. So we will see a lot of progress but to achieve human level efficiency and capabilities it will take longer. [...] Even LLM does not have human brain efficiency. Human brain operates on 30 watts. So we are far from that.

provides context · 2

Highlight slides
Spatial Intelligence: The Next Frontier✦ from: Spatial intelligence—the ability to generate, understand, reason with, and interact with physical or virtual spaces—is the next frontier of AI, and building large world models is the means to achieve it.World Labs: A Frontier Model Lab✦ from: Spatial intelligence—the ability to generate, understand, reason with, and interact with physical or virtual spaces—is the next frontier of AI, and building large world models is the means to achieve it.The Data Bottleneck in General-Purpose Robotics✦ from: The unique bottleneck in general-purpose robotics is the lack of data for training and evaluation, which can be solved by a 'real-to-sim-to-real' pipeline that aligns digital worlds with physical environments so that simulation data replaces real-world data at scale.Scenix: Real-to-Sim-to-Real Pipeline✦ from: The unique bottleneck in general-purpose robotics is the lack of data for training and evaluation, which can be solved by a 'real-to-sim-to-real' pipeline that aligns digital worlds with physical environments so that simulation data replaces real-world data at scale.Human-Level Robotics Efficiency: A Long Road Ahead✦ from: Human-level power efficiency in robotics will take a very long time to achieve because a working robot is always a system problem integrating hardware, software, and countless details, and we should take a measured, realistic approach rather than over-promising.The Efficiency Gap: LLM vs. Human Brain✦ from: Human-level power efficiency in robotics will take a very long time to achieve because a working robot is always a system problem integrating hardware, software, and countless details, and we should take a measured, realistic approach rather than over-promising.
Related episodes