ATRIUMsearch → argument graph
PredictionVideo · 11:04 — 12:27

A foundation model for robotics will be an omni-model that takes multimodal input—including actions as input (a forward simulator predicting environment changes) or actions as output (a policy model)—and serves as a backbone fine-tuned for specific robotic applications.

Fei-Fei Li and Yunzhu Li describe the foundation model for robotics as an omni-model that processes actions as inputs or outputs, acting as both a simulator and a policy backbone. ✦ AI generated

Fei-Fei Li and Yunzhu Li · a16z Podcast · 2026-07-28 · original ↗

starts at this moment · 11:04

Elicited by

So can we expect a foundation model for robotics from World Labs?

World Lab is building a foundation model. As you know, we're building a base model and as the technology has been evolving, some of the most exciting base models are omni-models, right? They take multimodal input, they have multimodal outputs. And what is a foundation model for robotics? It's very likely going to involve actions. It's very likely going to involve the output of actions in addition to the state of the world and we're definitely not ruling this out. [...] For the foundation models it's essentially needs to be a multimodal model. So it has to take into account frame, text, image, depth and different kind of modalities and action is a very important parts of that modalities. So if you think about using actions as inputs that is essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action. When the action is output this is essentially a policy model that is trying to predict given a specific goal what should be the action you take in the real environment to get you closer to that goal. So this kind of omni models actually can benefit a lot and provide huge amount of values for the robotics communities and this can also acts as a backbone for you to fine-tune into specific robotic applications.

verbatim transcript · starts at 11:04

Transcript · around this moment

10:46opportunities of leveraging like a marble and other like capabilities as world labs in order to do very efficient reconstructions and modeling of the environments. >> So can we expect a foundation model for robotics from world labs? World Lab is building a foundation model. As you know, Martin, we're building a base model and uh as the technology has been evolving, some of the most exciting base models are omniodels, right? They take they

11:16take multimodal input, they have multimodal outputs. And uh what is a foundation model for robotics? Uh it's very likely going to involve actions. M >> it's very likely going to involve the output of actions in addition to the state of the world and we're definitely not ruling this out. >> Yeah. Great. >> So for example for the foundation models it's essentially needs to be a multimodal model. So it has to take into

11:44account frame text image deps and different kind of modalities and action is a very very important parts of that modalities. So if you think about frame actions as inputs that is essentially a forward simulator that is going to predict how the environment is going to change when you apply a specific action. When the action is output this is essentially a policy model that is trying to predict given a specific goal

12:07like what should be the action you take in the real environment to get you closer to that goal. So this kind of omni models actually can benefit a lot and actually provide huge amount of values for the robotics communities in trying to understand how to model the environments and at the same time how to act in the environments and this can also acts as a backbone for you to

12:27fine-tune into specific robotic applications to making sure it's really leave up to the reliability and efficiency that expected by the clients. you know if uh Yun if you don't uh if you don't mind a kind of a lay investor question I see a lot of robotics companies and a very popular approach right now for the robotics companies that come in is like we'll use a video model you know and like you know that's

12:50the the predominant method where this is you know 3D and simulation it's a very different approach and so maybe you could contrast you know this popular approach of just using video only versus kind of what the ambition here is >> yeah so in order to create words where the robot can learn the words as I mentioned is to capture the essential structure of the problem and one of the

Around this claim