Scaling laws in large language models can be quantitatively predicted from two measurable linguistic quantities: the power-law decay rate of token correlations with distance and the power-law entropy reduction as context length increases.
Wyart's team proposed a theory that explains the empirical scaling laws driving massive AI investment: the training loss exponent can be derived from two measurable properties of natural language — how quickly word-word correlations decay with distance and how quickly entropy drops as more context is provided — both following power laws. ✦ AI generated
Matthieu Wyart · Machine Learning Street Talk · 2026-08-10 · original ↗
starts at this moment · 112:27
“About a year ago you had a paper about scaling laws as well. Tell me about that.”
the theory predicts that there is a simple recipe to for those exponents to extract those and and essentially this is saying again what's underlying it is the fact that if you give me more data I can learn more abstract concept and that's longer range leads to longer range correlation but at the end of the day the two quantities you need to measure is one the fact that words or tokens are correlated and that this correlation it's it was well known before that this correlation decreases as a polo of the distance between those two words. From that you can measure exponents and they depend on the language you look at as your data set. You can measure them and then there is another key quantity we argue which is related to the entropy of text. So entropy of text has been discussed already by Shannon in the 50s. It's beautiful question. So essentially it's the entropy is a log of the number of possible words that you would have at one location in average. What we argue is very important to look at and that we could finally measure with LLMs or other architecture and we find consistent result is what is the entropy left after a sentence of n token. If you see n token the more token you see the least possibility you have there. What is the entropy of that? And in the toy model of it's a polo and in real life it's also a polo that is found. And so essentially what we argue is that with those two exponents you can combine them in a way that we specify to get the training curve exponent of LLM's acting on those natural languages and it works very well.
verbatim transcript · starts at 112:27
(01:11:46) you know maybe nuclear plants. [laughter] So it has a huge technological impact but it's a bit embarrassing for us theories that essentially it's not understood at all. You know those scaling laws quantitatively they have exponents in them. For example describing how well you perform better if you multiply the number of data by 10 and there was very limited understanding uh on that question. Um and so yeah so
(01:12:16) just a few months back with Franchesco Kagneta uh Alan Ravventos and soya Ganguli we proposed a theory for this problem inspired by those synthetic world I told you about but uh detaching sort of essence of the lesson we learned from those models to really make quantitative prediction for natural languages and essentially we the theory predicts that there is a simple recipe to for those exponents to extract those
(01:12:47) and uh and essentially this is saying again what's underlying it is the fact that if you give me more data I can learn more abstract concept and that's longer range leads to longer range correlation but at the end of the day the two quantities you need to measure is one the fact that words or tokens are correlated and that this correlation it's it was well known before that this
(01:13:12) correlation decreases as a polo of the distance between those two words. From that you can measure exponents and they depend on the language you look at as your data set. You can measure them and then there is another key quantity we argue which is related to the entropy of text. So entropy of text has been discussed already by Shannon in the 50s. It's beautiful question. So essentially it's