Deep architectures have an implicit bias that enables them to learn hierarchical generative structure from exponentially fewer examples than Chomsky's poverty of stimulus argument predicts is possible.
Wyart's group showed that while shallow networks memorize and fail to learn grammar, deep architectures leverage their hierarchical structure as an implicit bias, learning to generate novel grammatical sentences from only polynomial (not exponential) amounts of training data — directly contradicting Chomsky's poverty of stimulus argument. ✦ AI generated
Matthieu Wyart · Machine Learning Street Talk · 2026-08-10 · original ↗
starts at this moment · 28:35
“is missing is the acquisition of abstractions. Now, you've done some amazing work on this, but certainly when I use language models, what's abundantly clear to me, even when we do hill climbing and we we solve these mathematical problems, they traverse the spaghetti monster and they get the right answer, but for the wrong reasons”
Chumsky gave this uh poverty of stimulus argument uh you know arguing that it was actually impossible to learn to become creative from example and essentially okay this would be a very crude way of of summarizing his argument but I described the fact that you have this sort of generative treel like rich contextf free grammarss you know assuming that that and really capturing the fact that the the world as a yarchy of abstract concepts. Um and but you have other possible generative grammar. Some are much simpler. One is called regular grammar. So this will be a carature but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chumsky's argument is to say that well even if you give me 1 million sentences okay I can fit those you know those sentences by a contextf free grammar but I can also fit them by a much simpler I mean a regular grammar simpler in it classification but to fit those sentences it would have to be awfully complicated you know many many many rules... So in our idealized world where the true world is can machine learn to be creative or not and what we find is that if you have a shallow network what Shamsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those course grain variables. Thisarchical architecture leads very easily to some iterative calculation. And so what we what we found is that uh indeed you can learn to be creative by having being exposed to very small number of sentences. So let's say in our model if D is an is you know the size of the sentence the number of sentences is huge with D is exponential in D but the number of sentences you need to see to be creative is only polomial in D. So these models are really a counter example to uh to his argument.
verbatim transcript · starts at 28:35
28:35with Shsky. So, so you know um so so there is a question creativity I will use this term in a very narrow sense of being able to generate new sentences that satisfy hard constraint syntactic rules that you know the child would never have heard before. and um and Chsky gave this uh poverty of stimulus argument uh you know arguing that it was actually impossible to learn to become creative
29:06from example and essentially okay this would be a very crude way of of summarizing his argument but I described the fact that you have this sort of generative treel like rich contextf free grammarss you know assuming that that and really capturing the fact that the the world as a yarchy of abstract concepts. Um and but you have other possible generative grammar. Some are much simpler. One is called regular
29:36grammar. So this will be a carature but essentially the idea that maybe a group of words will fix the probability of the next word essentially. And Chumsky's argument is to say that well even if you give me 1 million sentences okay I can fit those you know those sentences by a contextf free grammar but I can also fit them by a much simpler I mean a regular
30:00grammar simpler in it classification but to fit those sentences it would have to be awfully complicated you know many many many rules and uh and that you know nativism and operism big debates on this question and again we we we felt like we want to address those question as physicists. So in our idealized world where the true world is can machine learn to be creative or not
30:32and what we find is that if you have a shallow network what Shamsky worried about is completely true. You learn some you don't learn this sort of interesting generative grammar. You essentially memorize and you can't do anything. But if you have a deep architecture, there's a huge implicit bias to build those course grain variables. Thisarchical architecture leads very easily to some iterative calculation. And so what we what we found is that uh indeed
31:11you can learn to be creative by having being exposed to very small number of sentences. So let's say in our model if D is an is you know the size of the sentence the number of sentences is huge with D is exponential in D but the number of sentences you need to see to be creative is only polomial in D. So these models are really a counter
31:36example to uh to his argument and and ultimately it come from the fact that you know machines have strong implicit bias. they are not comparing equally in an equal fashion all hypothesis. Uh and if you're deep you learn so so uh so that's a counter example to that. So so in some sense I would argue that in terms of what needs to be innate if you have deep architecture that's
32:07[snorts] does lots of that. This is not to say that, you know, so I'm I'm arguing against uh Shamsky's argument. It doesn't mean that what he inferred is incorrect, right? It's not because I I think an argument is incorrect, that the statement is incorrect. I don't want to imply that our brain is just a deep net and that there are not, you know, much smarter mechanisms to learn, you know,
- ·Chomsky argued: learning creativity from examples is impossible
- ·Shallow networks confirm this — memorize, fail to learn grammar
- ·Deep architectures have implicit bias for hierarchical structure
- ·This bias enables learning creativity from very few examples
- ·Possible sentences grow exponentially with sentence length D
- ·Deep networks learn generative grammar from only polynomial(D) examples
- ·Direct counterexample to poverty of stimulus argument