ATRIUMsearch → argument graph
MechanismVideo · 39:24 — 40:54

Video generation models are evolving into 'world models' that predict not just how the world looks and sounds but what actions are possible in it, directly analogous to how LLMs predict the next token from internet-scale pretraining.

Zeve explains that video models are becoming 'world models' predicting the next moment — appearance, sound, and possible actions — drawing a direct analogy to next-token prediction in LLMs, and cites Nvidia's paper showing robot joint states can be encoded directly into video tokens. ✦ AI generated

Zeve Farman · The Cognitive Revolution · 2026-07-08 · original ↗

starts at this moment · 39:24

Elicited by

Zeve, just give us a brief introduction to what Litrix LTX has been doing for the last three months, what what has what has taken off?

there's like a growing realization that uh what started as video models is becoming a backbone of what we call now like world models. And I think like the best way to explain why this is so powerful is to use the analogy to LLM's right like in the end of the day at their core LLMs are still predicting the next talk in the next word.

verbatim transcript · starts at 39:24

Transcript · around this moment

39:24think there's like a growing realization that uh what started as video models is becoming a backbone of what we call now like world models. And I think like the best way to explain why this is so powerful is to use the analogy to LLM's right like in the end of the day at their core LLMs are still predicting the next talk in the next word. And when we

39:48do the pre-training at the scale of the internet it allows us to create models that do textual reasoning incredibly well. And uh the emerging world models they're kind of doing the same right like giving some kind of boundary condition some kind of history some kind of uh constraints they predict the next moment okay and the moment includes how the world appears how it like sounds and what kind of action we can do. I think

40:20the action part is the most maybe surprising one and like roughly I would say like a quarter ago maybe a bit more in video showed in their um dream zero paper that it's fairly easy to add to video talking some kind of encoding of the joints of the robot and then basically completely ditch the VA paradigm that was reigning supreme before it. So I think that's like one of the big

40:52surprises and for us realizing that was like this big moment that validated something that we always strive for is to create an extremely efficient models because I think once you start to realize that the robot will need to create like this simulation 30 times a second you just like realize the amount of tokens that is going to be burned for these simulations. So I think that was

41:17like one of the maybe exciting validations of the overall thesis in terms of architecture like a bunch of things that we can I don't know discuss in depth. We're planning to release soon our mixture of expert architecture besides the dense models that we're already releasing. I think we finally were able to crack variable tokens architecture which also exciting and kind of teaches the model to invest more

Around this claim