Video generation models are becoming the backbone of 'world models' that predict the next moment of world state — how it looks, sounds, and what actions are possible — the same way LLMs predict the next token from internet-scale pre-training.
Zeve Farman explains LTX's thesis that video models, by predicting the next 'moment' (appearance, sound, and possible actions), are becoming world models in the same way LLMs became general reasoners from next-token prediction at scale. ✦ AI generated
Zeve Farman · The Cognitive Revolution · 2026-07-08 · original ↗
starts at this moment · 39:24
“Zeve, just give us a brief introduction to what Litrix LTX has been doing for the last three months, what what has what has taken off?”
at their core LLMs are still predicting the next talk in the next word. And when we do the pre-training at the scale of the internet it allows us to create models that do textual reasoning incredibly well. And uh the emerging world models they're kind of doing the same right like giving some kind of boundary condition some kind of history some kind of uh constraints they predict the next moment
verbatim transcript · starts at 39:24
39:24think there's like a growing realization that uh what started as video models is becoming a backbone of what we call now like world models. And I think like the best way to explain why this is so powerful is to use the analogy to LLM's right like in the end of the day at their core LLMs are still predicting the next talk in the next word. And when we
39:48do the pre-training at the scale of the internet it allows us to create models that do textual reasoning incredibly well. And uh the emerging world models they're kind of doing the same right like giving some kind of boundary condition some kind of history some kind of uh constraints they predict the next moment okay and the moment includes how the world appears how it like sounds and what kind of action we can do. I think
40:20the action part is the most maybe surprising one and like roughly I would say like a quarter ago maybe a bit more in video showed in their um dream zero paper that it's fairly easy to add to video talking some kind of encoding of the joints of the robot and then basically completely ditch the VA paradigm that was reigning supreme before it. So I think that's like one of the big
40:52surprises and for us realizing that was like this big moment that validated something that we always strive for is to create an extremely efficient models because I think once you start to realize that the robot will need to create like this simulation 30 times a second you just like realize the amount of tokens that is going to be burned for these simulations. So I think that was
41:17like one of the maybe exciting validations of the overall thesis in terms of architecture like a bunch of things that we can I don't know discuss in depth. We're planning to release soon our mixture of expert architecture besides the dense models that we're already releasing. I think we finally were able to crack variable tokens architecture which also exciting and kind of teaches the model to invest more
- ·LLMs predict next word from internet-scale pre-training
- ·This yields strong general textual reasoning
- ·Video models now predict the next 'moment' similarly
- ·Moment includes appearance, sound, and possible actions
- ·Given history and constraints, models predict what comes next
- ·LLMs: next-word prediction → general reasoners
- ·Video models: next-moment prediction → world models