The same generative model that makes images and videos can be extended with action prediction so it can ultimately be deployed as the controlling brain on a physical robot.
Rombach describes Black Forest Labs' shift toward multimodal models that combine image, video, and audio generation with action prediction, enabling deployment on real-world robots. ✦ AI generated
Robin Rombach · All-In Podcast · 2026-07-10 · original ↗
starts at this moment · 43:17
We are now like entering a new paradigm which is combining that with something that's called action prediction such that you can actually use the same model to make images, to make videos, to make audio, and to predict actions — which means you can ultimately deploy it on a robot in the real world.
verbatim transcript · starts at 43:17
43:17can actually use the same model to make images to make videos to make audio and to predict actions which means you can ultimately deploy it on a robot in the real world. >> Wow. So from the image to the video, the audio and then eventually the real world with robotics and a real world model because if you can make the image you and you can train the model that means
43:43by default you understand the world. In order to make a video of the world you have to understand the world. Yeah. [snorts] >> And the objects in >> I I think that's Yeah. I think that's like a a really good like way to think about it. Yeah, it's like it's like an intuitive way uh to interact with the world, right? Like I I would say there's like these like
44:03complimentary forms of intelligence ultimately. There's like intuitive intelligence and then there's like a deep reasoning layer. Now ultimately you need for like a kind of like complete form you need both um and you need them to interact and I think like we've been approaching it more from like the intuitive side um images is like a very natural way to approach this whole field because it's not as computationally
44:25intensive as let's say video right [snorts] but now yeah I think like we're combining it it's converging into like a a multimodal model and yeah we see like exactly like pre-training on videos gives like implicit understanding of the physics of interactions with the real world and then you can get stuff like action prediction like robotics out of the same model. And with these models and the training
44:51there kind of um been a limitation in creating videos and creating images where the criticism of generative AI is it's a bit of a slot machine. I give a prompt, it gives me something back. But how did it come up with that the training data, but you know, maybe I want a different style? Maybe I want u a different color. Maybe I want a different uh you know, aesthetic.