Foundation models like AlphaFold are systematically biased and overconfident precisely on the novel, edge-of-knowledge questions scientists actually care about, but merging a small amount of ground-truth data via 'prediction powered inference' can correct the error bars while retaining statistical power.
Jordan describes empirical work showing AlphaFold's confidence intervals are narrow but wrong for understudied phenomena like quantum fluctuations tied to phosphorylation, and explains 'prediction powered inference' as a fix that blends foundation-model output with small amounts of ground truth. ✦ AI generated
Michael I. Jordan · Machine Learning Street Talk · 2026-05-20 · original ↗
starts at this moment · 20:25
“you did some analysis on those 200 million predicted proteins and you found they were very good but there was something missing but you could robustify them.”
What if I add a little bit of ground truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth. And the answer is yeah, there's a methodology. We've developed something called prediction powered inference that does exactly that.
verbatim transcript · starts at 20:25
20:05the past and it's hard to crystallize and so not many examples means it's quite possible alpha won't give out a great answer but it won't tell you that it doesn't give you out air bars and it doesn't spec but specifically on the question you're asking that's where I want the air bars and it didn't know about that question when it was you know built and designed Okay. All right. So
20:25now I have a good statistical question. What if I add a little bit of ground truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth. And the answer is yeah, there's a methodology. We've developed something called prediction powered inference that does exactly that. And so it'll cover the truth just like in a classical statistical setting,
20:45but it's using this rather highly biased architecture. And it's now it's not biased overall. In fact, its accuracy is high overall. But for the question I'm asking, it might be very biased. And that's going to happen a lot in science because scientists are rarely interested in just studying the past over again. They're interested in brand new things on the edge of knowledge. And that's where specifically these foundation
21:04models will be most poor and most highly biased. So there needs to be around any foundation model the ability to maybe collect a bit of ground truth data to merge it in with some procedure like this and then to give out a more trustable answer. That's all not science fiction. that's what can be done and what really needs to be done and I'm sure the AlphaFold people are on board
21:26with that that they would not find that weird or surprising. Um, but a lot of other people out there talk about bias and all that and they either don't worry about it. I say it'll go away if we have enough data or they just critique the architectures and critique the outputs but they have no scientific I you know method in in mind that'll help us go forward. So that's kind of the state
21:47we're in. I challenged John a little bit about the extent to which AlphaFold understands and he was basically allergic to the word understands. >> We are not trying to tell you everything. We are not a model of the entire cell. These machines let us predict. They let us control. We have to derive our own understanding at this moment. Right? We can experiment now on the artifact. We can look at the
22:15200 million predicted structures. not just the 200,000 experimental structures in order to help us understand, but it doesn't do the act of understanding for us. It does the act of predict and maybe control. >> Why why should AlphaFold understand? >> Well, what would it mean to I mean, he was he was sketching it out to me. He kind of said that this this is a weird alien artifact and it's not like it's