The loss function of an overparameterized neural network undergoes a phase transition analogous to the jamming transition in sand: underparameterized networks have a rough energy landscape with many metastable states, while sufficiently many parameters allow the system to flow into flat zero-energy valleys.
Wyart connects his work on complex systems to machine learning: overparameterized networks' loss landscapes have the same phase transition as tilted sand flowing. Underparameterized models get stuck in metastable states, but with enough parameters the landscape has flat zero-energy valleys; this is what others independently called double descent but physicists see as a jamming transition. ✦ AI generated
Matthieu Wyart · Machine Learning Street Talk · 2026-08-10 · original ↗
starts at this moment · 4:40
what we discover is actually that this landscape has exactly the same phase transition as send it means that when you actually underparameterized when you don't have enough parameters you have a rough landscape with many metastable state and if you train your machine and you train it many times it will end up in different position where it's actually stuck but if you have enough parameters then suddenly the system can flow I in your landscape has many flat valleys and you can which have essentially zero energy. So there is really a close analogy we disco we discovered that like 9 years ago at the same time uh others find a very similar I mean the same phenomenon and called it double descent. So now this name has stuck but this double descent this peak of the double descent is really for physicists a jamming transition.
verbatim transcript · starts at 4:40
4:22some point it's going to flow. It it means that the energy landscape was rough and you are in a metastable state but you tilted this energy landscape you had a phase transition and then the entire system flow although it's very dense particle managed to avoid each other and so I've been very interested in you know understanding geometrically those questions. Uh but then like nine years ago uh I'm I'm a go player a poor
4:47go player but I enjoy playing and uh and I was mesmerized by you know um Alph Go and so on and so I started to think about machine learning and I started to think of it as a complex system and it is because when you train a machine you know you build a function that is low if you fit well your data it's called the loss function or cost function and so we
5:11are very intrigued by what is the geometry of this landscape and so what we discover is actually that this landscape has exactly the same phase transition as send it means that when you actually underparameterized when you don't have enough parameters you have a rough landscape with many metastable state and if you train your machine and you train it many times it will end up in different position where it's
5:34actually stuck but if you have enough parameters then suddenly the system can flow I in your landscape has many flat valleys and you can which have essentially zero energy. So there is really a close analogy we disco we discovered that like 9 years ago at the same time uh others find a very similar I mean the same phenomenon and called it double descent. So now this name has
6:02stuck but this double descent this peak of the double descent is really for physicists a jamming transition. So yeah, so so that brought me to machine learning and just maybe to uh to finish with that I mean we uh in the last four years we've been very much interested in another landscape that I think is even more interesting. It's a landscape of data. Uh so if you think
6:26about you know an image let's call X an image it's a vector you could ask what is the density of those images row of X and this question relates to what is the structure of the world and we think it's key to actually understand how a machine work >> does it make sense to talk about because obviously you know you're a physicist and you're applying this lens of