The loss function of an overparameterized neural network undergoes a phase transition analogous to the jamming transition in sand: underparameterized networks have a rough energy landscape with many metastable states, while sufficiently many parameters allow the system to flow into flat zero-energy valleys.
Wyart connects his work on complex systems to machine learning: overparameterized networks' loss landscapes have the same phase transition as tilted sand flowing. Underparameterized models get stuck in metastable states, but with enough parameters the landscape has flat zero-energy valleys; this is what others independently called double descent but physicists see as a jamming transition.
transcript
Matthieu Wyart: what we discover is actually that this landscape has exactly the same phase transition as send it means that when you actually underparameterized when you don't have enough parameters you have a rough landscape with many metastable state and if you train your machine and you train it many times it will end up in different position where it's actually stuck but if you have enough parameters then suddenly the system can flow I in your landscape has many flat valleys and you can which have essentially zero energy. So there is really a close analogy we disco we discovered that like 9 years ago at the same time uh others find a very similar I mean the same phenomenon and called it double descent. So now this name has stuck but this double descent this peak of the double descent is really for physicists a jamming transition.
gives example · 1supports · 1