Video generation models are evolving into 'world models' that function analogously to LLMs — instead of predicting the next word, they predict the next 'moment,' including how the world looks, sounds, and what actions are possible.
Zeve explains that video models are becoming general-purpose 'world models,' using the LLM next-token-prediction analogy to describe how they now predict entire future moments including sound and possible actions (relevant to robotics). ✦ AI generated
Zeve Farman · The Cognitive Revolution · 2026-07-08 · original ↗
starts at this moment · 39:24
“Zeve, just give us a brief introduction to what Litrix LTX has been doing for the last three months, what what has what has taken off?”
there's like a growing realization that uh what started as video models is becoming a backbone of what we call now like world models. And I think like the best way to explain why this is so powerful is to use the analogy to LLM's right like in the end of the day at their core LLMs are still predicting the next talk in the next word.
verbatim transcript · starts at 39:24
39:24think there's like a growing realization that uh what started as video models is becoming a backbone of what we call now like world models. And I think like the best way to explain why this is so powerful is to use the analogy to LLM's right like in the end of the day at their core LLMs are still predicting the next talk in the next word. And when we
39:48do the pre-training at the scale of the internet it allows us to create models that do textual reasoning incredibly well. And uh the emerging world models they're kind of doing the same right like giving some kind of boundary condition some kind of history some kind of uh constraints they predict the next moment okay and the moment includes how the world appears how it like sounds and what kind of action we can do. I think
40:20the action part is the most maybe surprising one and like roughly I would say like a quarter ago maybe a bit more in video showed in their um dream zero paper that it's fairly easy to add to video talking some kind of encoding of the joints of the robot and then basically completely ditch the VA paradigm that was reigning supreme before it. So I think that's like one of the big
40:52surprises and for us realizing that was like this big moment that validated something that we always strive for is to create an extremely efficient models because I think once you start to realize that the robot will need to create like this simulation 30 times a second you just like realize the amount of tokens that is going to be burned for these simulations. So I think that was
41:17like one of the maybe exciting validations of the overall thesis in terms of architecture like a bunch of things that we can I don't know discuss in depth. We're planning to release soon our mixture of expert architecture besides the dense models that we're already releasing. I think we finally were able to crack variable tokens architecture which also exciting and kind of teaches the model to invest more
- ·Video models evolving into general-purpose 'world models'
- ·Analogy: LLMs predict the next word, at their core
- ·World models predict the next 'moment' instead
- ·Covers how the world looks, sounds, and possible actions
- ·Predicting possible actions ties directly to robotics
- ·Same next-token logic, applied to physical world moments
- ·Positions video models as a new AI backbone, not just media tools