Reinforcement learning will move earlier and earlier into pre-training, and next-token prediction alone is not extracting enough from the web.
Eiso believes RL will shift earlier into the training process, and that current next-token prediction is leaving most of the web's learning potential untapped. He views distillation and synthetic environments as 'drugs' the industry is addicted to, and thinks the real gains will come from teaching models to think earlier in training. ✦ AI generated
Eiso Kant · Latent Space · 2026-07-23 · original ↗
plays this moment only · 43:43 — 45:13
I have a not commonly held opinion that reinforcement learning will move earlier and earlier into training. We've been spending a couple of years really doing research on how can we turn the web into not just next token prediction, but into a way to teach the model to think earlier in its training. And I think there's a huge amount of gold to be found there.
verbatim transcript · starts at 43:43
Eiso Kant [00:43:15]: Training is not done. I mean, look, there’s a part of training of just dealing with skill, right? Every new order of magnitude of model skill, you are going to get new things you gotta solve for. That’- but those are ultimately, engineering challenges.
Eiso Kant [00:43:31]: I have a, I would say, a not commonly held opinion that reinforcement learning Will move earlier and earlier into training.
Vibhu [00:43:42]: Yeah, training.
Eiso Kant [00:43:44]: Not even training. Like training today, right, is, like if you look at - So we’ve been working on this for years already. and I think the best-- I think the first time we saw it out in public was the DeepSeek Zero paper. this is a year and a half ago, I think, if I recall correctly. where, you can Very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning. and so the question that I have is like, we have this- we have the dataset that’s the web. and the web, I think we could arguably say probably has The totality of humanity’s knowledge somewhere encoded in different places. It’s a huge variance degree of quality, from garbage data, and like once you look at training data, you really get humbled of like what the web is, to like, the most greatest scientific papers and best blog posts and like, best transcripts and whatnot.