Foundation models like AlphaFold are highly accurate overall but give badly overconfident, biased predictions on edge-of-knowledge questions with little training data; merging a small amount of ground-truth data via 'prediction-powered inference' can correct the error bars while preserving statistical power.
Using AlphaFold as an example, Jordan describes finding that its predictions on rare, edge-of-knowledge questions (quantum fluctuations and phosphorylation) were narrowly overconfident and far from the true value, and describes 'prediction-powered inference,' a method his group developed to fix this by blending in a bit of ground-truth data. ✦ AI generated
Michael Jordan · Machine Learning Street Talk · 2026-05-20 · original ↗
starts at this moment · 19:42
“I interviewed John Jumper last week at Google and you did some analysis on those 200 million predicted proteins and you found they were very good but there was something missing but you could robustify them.”
What if I add a little bit of ground truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth. And the answer is yeah, there's a methodology. We've developed something called prediction powered inference that does exactly that. And so it'll cover the truth just like in a classical statistical setting, but it's using this rather highly biased architecture.
verbatim transcript · starts at 19:42
19:42on that statistic of that 2x two table was extremely narrow and way far from the truth the true value of the uh the the gold standard value. And we found this in domain after domain you know. So why is that? Well, what's happening there is that there's probably not many um examples in the training set of proteins with quantum fluctuation because it's not been that studied in
20:05the past and it's hard to crystallize and so not many examples means it's quite possible alpha won't give out a great answer but it won't tell you that it doesn't give you out air bars and it doesn't spec but specifically on the question you're asking that's where I want the air bars and it didn't know about that question when it was you know built and designed Okay. All right. So
20:25now I have a good statistical question. What if I add a little bit of ground truth data to the 200 million? Can I shift the error bar so it stays somewhat narrow? So I have high power, but it covers the truth. And the answer is yeah, there's a methodology. We've developed something called prediction powered inference that does exactly that. And so it'll cover the truth just like in a classical statistical setting,
20:45but it's using this rather highly biased architecture. And it's now it's not biased overall. In fact, its accuracy is high overall. But for the question I'm asking, it might be very biased. And that's going to happen a lot in science because scientists are rarely interested in just studying the past over again. They're interested in brand new things on the edge of knowledge. And that's where specifically these foundation
21:04models will be most poor and most highly biased. So there needs to be around any foundation model the ability to maybe collect a bit of ground truth data to merge it in with some procedure like this and then to give out a more trustable answer. That's all not science fiction. that's what can be done and what really needs to be done and I'm sure the AlphaFold people are on board
21:26with that that they would not find that weird or surprising. Um, but a lot of other people out there talk about bias and all that and they either don't worry about it. I say it'll go away if we have enough data or they just critique the architectures and critique the outputs but they have no scientific I you know method in in mind that'll help us go forward. So that's kind of the state