ATRIUMsearch → argument graph
MechanismVideo · 44:42 — 45:40

Concepts and abstractions in deep generative models emerge from statistics alone by grouping together configurations that predict similar contexts around them.

In LLMs and diffusion models, abstract concepts like 'street' or 'house' are not predefined but instead emerge when the network groups together diverse low-level configurations (passersby, cars, sidewalks) that all predict similar surrounding contexts. ✦ AI generated

Matthieu Wyart · Machine Learning Street Talk · 2026-08-10 · original ↗

starts at this moment · 44:42

if you think about uh LLMs or diffusion models, the way they build concepts, they emerge from statistics alone. Those abstraction they emerge they are there in the data they emerge and those concept emerge if you group together configuration that predict similar context around them. So maybe if you have a street typically you have houses nearby and maybe the houses have colors or edges and so you would predict color and edges and with that in those models at least you find that if you have enough data you can learn all the abstractions but as you get more and more abstract you have a problem because you're always trying to build those abstraction by saying how they are predictive but at a very low level and when you're very abstract you how you predict pixels or colors and so on is a super noisy signal. So essentially that's why in those models we find and we have empirical evidence and I'm happy to talk about empirical evidence that that the more abstract concept are you know the toughest to learn because essentially your signal as you get more and more abstract your signal gets diluted.

verbatim transcript · starts at 44:42

Transcript · around this moment

44:42concepts, they emerge from statistics alone. Those abstraction they emerge they are there in the data they emerge and those concept emerge if you group together configuration that predict similar context around them. So maybe if you have a street typically you have houses nearby and maybe the houses have colors or edges and so you would predict color and edges and with that in those models at least you find that if you have [snorts]

45:13enough data you can learn all the abstractions but as you get more and more abstract you have a problem because you're always trying to build those abstraction by saying how they are predictive but at a very low level and when you're very abstract you how you predict pixels or colors and so on is a super noisy signal. So essentially that's why in those models we find and we have empirical

45:40evidence and I'm happy to talk about empirical evidence that um that the more abstract concept are you know the toughest to learn because essentially your signal as you get more and more abstract your signal gets diluted. >> Yeah. >> So this is yes. So this would be the mechanism whereby we think you build those uh those latent variables or abstractions. >> Yeah, it it's such it's such a

46:05tantalizing idea and we'll we'll get to the the latent stuff just in a minute because that that's also a great a great thing to talk about. But um are you suggesting that there is some kind of natural factorization? So do you think that different networks perhaps with different architectures given the same data would almost converge towards the same logical factorization of the data this hierarchical factorization. So in

Around this claim