Chomsky's poverty of stimulus argument, that it is impossible to learn to become creative from examples alone, is refuted: deep architectures have a huge implicit bias to build coarse-grained hierarchical variables, so they can learn to be creative from polynomially many (not exponentially many) sentences.
In Wyart's synthetic tree-structured world, a shallow network does exactly what Chomsky predicted — it memorizes and cannot generalize. But deep architectures exhibit a strong implicit bias to construct coarse-grained hierarchical variables, learning to be creative from only polynomially many sentences. This is a counterexample showing that what must be 'innate' shrinks dramatically for a deep architecture. ✦ AI generated
Matthieu Wyart · Machine Learning Street Talk · 2026-08-10 · original ↗
starts at this moment · 16:27
“Just just to kind of play that back just so that everyone understands that the idea is that there is I mean we're talking about um grammar here but you know more more broadly we think that there are structured generative processes in the world.”
what we find is that if you have a shallow network what Shamsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those course grain variables. Thisarchical architecture leads very easily to some iterative calculation. And so what we what we found is that uh indeed you can learn to be creative by having being exposed to very small number of sentences. So let's say in our model if D is an is you know the size of the sentence the number of sentences is huge with D is exponential in D but the number of sentences you need to see to be creative is only polomial in D. So these models are really a counter example to uh to his argument and and ultimately it come from the fact that you know machines have strong implicit bias. they are not comparing equally in an equal fashion all hypothesis. Uh and if you're deep you learn so so uh so that's a counter example to that. So so in some sense I would argue that in terms of what needs to be innate if you have deep architecture that's lots of that.
verbatim transcript · starts at 16:27
16:33you need your theory to be and that depends the question you're asking >> and who who were your mentors who inspired you know what books did you read how did you kind of land on your current trajectory as a physicist >> well yeah that's a complex question for me because I have my two my two parents are physicist actually and and when I started to do my PhD I I tried to escape
16:56uh to escape them by going to econophysic. I mean doing more finance and economy and then already during my PhD I started to be fascinated again by physics and how sand flows and and things like that. Um and then when I was a postoc actually I was always mesmerized by our brain how we think uh and uh and I I tried at that time to uh you know I even spend one year in
17:22Jalia farm a place where you know a neuroscience institute and and I met lots of fantastic people and I learned a lot but I felt at that stage that a lot of the theory you know how you know you have many connected neurons what their dynamics a bit applied math and detach from really function and then the question of how do you learn intelligence or rules constraint language and so and so I was
17:47not I was not at the it was not at the level that I really wanted to operate and so I gave up I I went back to physics and then uh I mean I think it's essentially the development of the technology it's uh it's facing us so so I think the analogy I one analogy I I like is uh you know the industrial revolution when uh when
18:11actually um heat engine emerged before again one case where technology was first and uh and then you you had to understand you know what's behind them how efficient can they be the limit to their efficiency and so and so then caro actually a French physicist uh came up and wrote a you know beautiful text it reads like philosophy there's essentially no math and introducing concept like entropy and it was the
18:36beginning of thermodynamics. I mean very deep ideas coming from some uh technological fact. And here I think it's the same. I mean with those machines it's amazing. You look they are creative. You you give them a bunch of images and those diffusion models suddenly like a painter they you know they build new faces they compose new faces. How can it be? Or they create sentences that they have never heard
19:01before. And Nam Shamsky and others said that it would be extremely hard to do. uh they do it. So how why [laughter] so yeah it's being fascinated by questions. Yeah that's that's what drives me. Yeah I mean you know sorry to bring Chsky back again but he he said that LLMs are like bulldozers. You know he he says I love bulldozers. They're great for clearing the snow but they're
19:22not a contribution to science. And and he said you know he's I' I've got a theory. Anything goes right. You know it explores all the laws of nature. anything that can be and you know he says that when you've got a scientific theory you have to explain why are things this way why are things not that way but you were just saying you know when when we discovered the steam engine
- ·Shallow networks confirm Chomsky: memorize, cannot generalize
- ·Deep architectures build coarse-grained hierarchical variables
- ·Creative learning needs only polynomially many sentences