ATRIUMsearch → argument graph
MechanismVideo · 10:52 — 13:08

Mean squared error loss is fundamentally inadequate for generating functional neural network weights from a latent representation, because it reconstructs average weights well but loses the high-frequency information critical for network function.

Borth explains the 'blurry weights' problem: their autoencoder achieved very low reconstruction error, but when generated weights were loaded into a network, the network failed completely because MSE captures the average well but discards the precise high-frequency structure that makes a network functional. ✦ AI generated

Damian Borth · The TWIML AI Podcast · 2026-07-27 · original ↗

starts at this moment · 10:52

Elicited by

I can imagine you know lots of different directions including scaling up the models uh trying to get more insights out of the space like what was next

We we trained this out encoder. The mean split error was super low. We took the neural networks. We moved them in the forward pass. We reconstructed as said the loss is very low. We pluck the weights back to the neural network. Totally screwed up the entire new network. We're like [laughter] yeah it was like it was really like the mean spread is low and obviously now obviously right a mean square is an average. So we were very good at know reconstructing the average weights but the little things that make the difference of having this function working or not they were so important and there is analogy to pixels and images like when you had generative models for images the images were always blurry so people kind of tried hard and you know taming transformers for high resolution images changed the mean square error to perception laws and you know did some additional things on quantizing and the gun so we knew that reconstruction maybe a wrong loss and we tried to normalize and play around with the losses and you know thought about the behavior loss and then we were able to make those models little bit better but there's still a little bit of delta that is missing we kind of we generate at that time blurry weights right low frequency information high frequency information missing

verbatim transcript · starts at 10:52

Transcript · around this moment

10:52you know lots of different directions including scaling up the models uh trying to get more insights out of the space like what was next the the thing is for this type of research everything was very obvious it was like lying out and you just needed to do it I mean it's I never had this before right so you have an autoenccoder you take the encoder so you can predict

11:13discriminative downstream task like what's the accuracy what's you know generalization gap but we have the other thing called the decoder So can we sample from this space to generate neuronet networks? It's obvious, right? We we didn't have space in the first paper. So we needed one more year to that had a 22 paper published on generating neural networks. Hopefully those neural networks were then better than you know standard initializations.

11:39They were not as good as final or fully trained neural networks. So there was some there was some trouble that we had which was really really interesting. We we trained this out encoder. The mean split error was super low. We took the neural networks. We moved them in the forward pass. We reconstructed as said the loss is very low. We pluck the weights back to the neural network.

12:00Totally screwed up the entire new network. We're like [laughter] yeah it was like it was really like the mean spread is low and obviously now obviously right a mean square is an average. So we were very good at know reconstructing the average weights but the little things that make the difference of having this function working or not they were so important and there is analogy to pixels and

12:25images like when you had generative models for images the images were always blurry so people kind of tried hard and you know taming transformers for high resolution images changed the mean square error to perception laws and you know did some additional things on quantizing and the gun so we knew that reconstruction maybe a wrong loss and we tried to normalize and play around with the losses and you know thought about

12:48the behavior loss and then we were able to make those models little bit better but there's still a little bit of delta that is missing we kind of we generate at that time blurry weights right low frequency information high frequency information missing and we know we were happy because again you know people were kind to us and said like it's toy examples but you know it's interesting

13:08we generate that the numbers are good but then I said okay we cannot be three times lucky so we to work really hard to scale that up, right? I mean, you know, you're lucky twice, but you know, three times, you know, your karma is gone for the next years. So, we we we then really and I'm very thankful to, you know, Constantine Sherhood who was part of

13:25that initial phase and he really worked hard and we had this idea of instead of reconstructing the entire model, think about the model parameters as a sequence that you window and then reconstruct the windows. And therefore we would kind of deattach the sequence left of the original model to the auto autoenccoder one. And this was interesting because suddenly we could go to restn nets and beyond. And this leds then to you know

Around this claim
This moment responds to