ATRIUMsearch → argument graph
AnecdoteAudio · 2:22 — 2:37

The goal of curing, preventing, and managing all disease by the end of the century was initially met with ridicule from scientists, but now that timeline may be too conservative due to the convergence of AI and biological data.

Mark Zuckerberg and Priscilla Chan recount how Nobel laureate scientists laughed at their initial ambition to cure all disease by 2100. A decade later, with AI advances, they believe that goal may have been too conservative. ✦ AI generated

Priscilla Chan · No Priors · 2026-06-10 · original ↗

plays this moment only · 2:22 — 2:37

Elicited by

Mark, Priscilla, thank you for doing this. Alex, congratulations on new missions. Thank you. guys made Biohub your primary philanthropic effort and then committed $500 million to this virtual biology initiative. Can you tell us a little bit about why do that and how did you go from we should fund this to this is like who we are?

So Biohub in its current form, we're super excited about. We feel like it's a really good fit for who we are and what we bring to the table and what we can achieve together. But this work started 10 years ago when we were thinking about how can we give back. And Mark wanted to build an organization that could cure, prevent, and manage all disease by the end of the century. And we had a series of hilarious meetings with scientists that like famous Nobel Prize winning scientists were just laughing at us. Is that, was that your starting line? We're just going to cure all disease. No, and to be clear, we don't think that we're going to be the ones curing the diseases. Our goal was always to build tools that could accelerate the whole scientific fields. That way the scientific field collectively could cure all the diseases. But still, people thought that by the end of the century was a stretch. Now I think it's like too conservative.

verbatim transcript · starts at 2:22

Transcript · around this moment

(00:00:00) We just want to give tools to the whole scientific community. (00:00:02) We want to understand how biology works. (00:00:05) I want to understand the genetics of this person. (00:00:07) I want to understand the risks they have to different illnesses. (00:00:10) My goal is to be able to treat the individual as an individual, understand the mechanisms, and be able to intervene. (00:00:17) We'll have a bigger impact by getting this in more scientists' hands quicker by doing it as open source projects instead. (00:00:23) It's not just like there's some factory somewhere that you can pay to produce the data. (00:00:28) You actually need to invent new, novel (00:00:30) scientific approaches, the theory isn't that we're going to cure the diseases, we're not. (00:00:34) It's that we want to help accelerate the pace of progress for the whole scientific field. (00:00:38) We folded over 1.1 billion proteins and predicted their structures, and we didn't design a model for antibodies. (00:00:44) We didn't design a model to be able to bind one particular target. (00:00:48) We just designed a model that could understand proteins. (00:00:50) If we could design A protein to actually change the physiology, then we can actually cure someone. (00:01:05) Today on No Priors, we're joined by Mark Zuckerberg, Priscilla Chan, and Alex Reeves. (00:01:09) We'll be talking about Biohub and all their various efforts to now start applying AI at scale to do world models of cells and different levels of interactions across biology. (00:01:19) Mark, Priscilla, thank you for doing this. (00:01:22) Alex, congratulations on new missions. (00:01:26) Thank you. (00:01:27) guys made Biohub your primary philanthropic effort and then committed $500 million to this virtual biology initiative. (00:01:35) Can you tell us a little bit about why do that and how did you go from we should fund this to this is like who we are? (00:01:42) So Biohub in its current form, we're super excited about. (00:01:46) We feel like it's a really good fit for who we are and what we bring to the table and what we can achieve together. (00:01:52) But this work started 10 years ago when we were thinking about how can we give back. (00:01:58) And Mark wanted to build an organization that could cure, prevent, and manage all disease by the end of the century. (00:02:05) And we had a series of hilarious meetings with scientists that (00:02:12) like famous Nobel Prize winning scientists were just laughing at us. (00:02:15) Is that, was that your starting line? (00:02:17) We're just going to cure all disease. (00:02:18) No, and to be clear, we don't think that we're going to be the ones curing the diseases. (00:02:22) Our goal was always to build tools that could accelerate the whole scientific fields. (00:02:27) That way the scientific field collectively could cure all the diseases. (00:02:31) But still, people thought that by the end of the century was a stretch. (00:02:34) Now I think it's like too conservative. (00:02:37) And so we kept being like, okay, well, we had these series of (00:02:41) funny, awkward, educational conversations where we're like, okay, but like, why? (00:02:45) Like, why do you think it's impossible? (00:02:48) And like, just being the person in the room is just like, oh, I don't know why. (00:02:53) You tell me. (00:02:54) Finally, we got people to like, they're like, fine, if you really must know. (00:02:58) And we're like, you know, we do. (00:02:59) It seems important. (00:03:02) It's, you know, they were like, well, we work in silos. (00:03:05) And when you publish, information doesn't get shared. (00:03:08) It gets locked up for long periods of time. (00:03:11) And we don't have tooling. (00:03:13) they gave the example of like, we build a great tool by 1 postdoc in a lab and it lives on their computer. (00:03:20) And when they graduate, the tool is gone. (00:03:22) And they just, it was, what we heard was very hard to build shared tools to move science faster. (00:03:30) build a shared knowledge base to quickly move science faster. (00:03:33) And that's sort of where we begin in thinking about, okay, like if those are the problems, like what can we contribute? (00:03:41) Yeah, I mean, so the original Biohub model was basically focused on long-term tool development by bringing together engineers and scientists across multiple universities. (00:03:52) to focus on long-term tool development. (00:03:54) And it basically, it like worked. (00:03:56) And we started off with CZI doing a number of different things. (00:04:01) And I think over time, we just felt like, okay, the science piece is really working. (00:04:05) And we just kept on investing more and more and more in it until now it is basically the primary and main thing that we're doing. (00:04:12) And we've expanded the original San Francisco Biohub to a handful now at this point. (00:04:18) There's New York, there's Chicago. (00:04:20) The real focus and the unifying (00:04:23) theme at this point is the virtual biology initiative around taking the unique data sets that are able to be generated in order to model, effectively starting with the smallest pieces of proteins, but then eventually cells and whole biological systems. (00:04:42) But that's kind of how we've evolved is, you know, this idea that we talk about around that some of this is an AI problem, (00:04:52) and you want to build a frontier AI lab, but you need to couple that with a frontier biology effort that can do the work of basically being able to understand and get the data that you need to actually be able to build these models. (00:05:08) Because unlike language models, there's just like a lot of data out there on the internet. (00:05:12) That's not really the case with biology. (00:05:14) I mean, there are obviously a bunch of different data sets that exist that academia and scientists have generated over the decades. (00:05:21) But (00:05:22) A lot of the stuff that I think we want to put into this, it doesn't exist, right? (00:05:25) It's like you want to be able to visualize things that people haven't been able to see before, which is why we're doing the imaging work. (00:05:30) You want to be able to record things that are going on inside the body, which is why we're doing the kind of cellular engineering work. (00:05:36) You want to be able to measure things like inflammation in ways that haven't been possible, which is why the Chicago Biohub is focused on building those kind of devices and being able to do that. (00:05:45) And that will fundamentally create (00:05:48) new types of data sets that will allow new types of models. (00:05:51) And I think it's just a very exciting thing that, I'm going back to what you're saying, if the scientific field, it primarily needs kind of tool development that now is going to empower scientists across the field to be able to do their work faster, that's what we think we can provide through this kind of long-term focus on tool development. (00:06:11) But I think there's a fun through line on (00:06:15) where we started and bringing us to our work to with that Alex is driving now is that our very first request for application RFA here was around single cell sequencing. (00:06:28) And we wanted to look at sort of like the RNA that is transcribed in individual cells. (00:06:34) And that was possible, but it was still pretty early on in understanding how different cells were expressing their DNA. (00:06:41) to the point where at the beginning, we were just funding methods, like getting people to describe how to do it so that others could share that methodology. (00:06:50) And then that became us funding the Human Cell Atlas, which is now one of the largest databases of single cell transcriptomes. (00:06:59) It was getting hard for scientists to annotate the data. (00:07:02) So we built Cell by Gene, which was like a very simple annotation tool that scientists could use to make use of that data. (00:07:08) Then a community came around Cell by Gene, built around Cell by Gene, and started contributing more and more data that we had nothing to do with sort of creating or funding or making happen in the world. (00:07:21) And now Cell by Gene is a corpus of knowledge that a lot of the transcriptomic-based models are based off of and is used regularly by the scientific community. (00:07:33) But still, there are always critiques. (00:07:35) Like, this is just stamp collecting. (00:07:37) Like, you're just gathering bits of knowledge, sorry, bits of data. (00:07:43) And we're not going to be able to pull scientific knowledge and wisdom and insights out of. (00:07:48) And we're like, well, we didn't have an answer for a while. (00:07:51) And then imagine our delight when large language models became a huge topic of conversation that could make sense (00:07:59) of large amounts of data. (00:08:01) And I just, for me, it was like, what if we could actually understand how biology worked, move it from a discovery-based science to an engineering-based science, where we could systematically understand how living beings, living cells worked and be able to understand why things go wrong. (00:08:21) And so when we saw that moment, we're like, this is it. (00:08:25) Something really big could happen here. (00:08:28) Alex, you were, you started at Meta Fair, but you were on the path to, you'd assemble the team at evolutionary scale and you raised venture and you were making progress in your models. (00:08:38) What was the pitch from Mark and Prasilia where you said like, that's actually the right way to go after the mission? (00:08:43) I think for me, it was really kind of the moment when I understood that, they really saw this as an integration of frontier AI and frontier biology. (00:08:54) And I think I had developed conviction that, this is really a new era of science that's just beginning, kind of what's going to be possible with artificial intelligence and (00:09:05) we're in the age of information theory at scale, and we have these systems that can basically kind of predict the next token, and they can, learn world models from that. (00:09:15) They can learn biology from the data. (00:09:17) And so, I think that it just, it was really clear that, to build kind of that next (00:09:24) that next kind of institution for the next era, you would really need to have frontier artificial intelligence. (00:09:30) You would have to have frontier biology. (00:09:32) You would need to start to put those things in feedback and really have models that are learning from the biology. (00:09:38) And I think, you know, it's just, and you'd need the right scale and the right people. (00:09:42) And so this just really felt, I think, like the way to do that. (00:09:46) There's A variety of different models that you all have been working on. (00:09:48) And I think it's kind of interesting because (00:09:51) Some of the earliest breakthroughs in biology were things like AlphaFold, where there was a Google model that showed that you could do protein folding at scale in a really interesting way that people didn't realize was very tractable. (00:10:00) And this was pre sort of the really big transformer waves that came later. (00:10:03) And then you're working on a variety of different things at different scale, right? (00:10:06) You're doing incremental molecular modeling and protein folding. (00:10:09) You're doing cell-based stuff. (00:10:11) You're thinking about interrogating larger scale systems in biology. (00:10:15) How well do you think that extends from sort of the micro to the macro? (00:10:17) You mentioned almost starting with building blocks and building up, but modeling cellular behavior is very different from modeling protein folding. (00:10:23) The data is very different. (00:10:24) The modeling is different. (00:10:25) I'm just curious, like, do you think it's all similar in terms of just data and you train stuff? (00:10:30) Or do you think it's actually, there's some differences in terms of how you actually have to deal with these systems? (00:10:36) I mean, there are probably some differences. (00:10:37) I mean, you can probably talk more to the specifics around this, but like, (00:10:41) I mean, I think each layer is going to end up being somewhat qualitatively different, right? (00:10:45) I mean, the, but you need to be able to understand the protein interactions in order to be able to understand how cells work. (00:10:50) So you can't just go straight to cells in a way without understanding the protein modeling. (00:10:55) And then if you're trying to understand something like the, you know, the way the immune system works or a bunch of cells interact together, then (00:11:02) it's tough to do that without first understanding cells. (00:11:05) I mean, you might be able to, at a very high level of abstraction, simulate a system. (00:11:09) But if you really want to understand how it's going to work, you kind of want to build the simulations at each level hierarchically. (00:11:14) So that's basically the approach that we're going through, starting with the building blocks and the protein. (00:11:20) But yeah, I mean, I think that there's going to be different types of data that you want to collect for each. (00:11:25) The modeling techniques, I think we'll see. (00:11:27) I mean, that'll all keep on advancing across the board. (00:11:29) But I do think that a big part of the strategy is this view that you need to build it up hierarchically. (00:11:35) And one of the things that's unique about us in the space is we were very intentional. (00:11:40) that the AI efforts and the wet lab efforts were a single effort. (00:11:46) And we've done a lot of work to bring them together. (00:11:48) And the really neat thing that we can do is really try to pull and gather data that helps us connect across sort of the hierarchy. (00:11:59) You know, you can look (00:12:00) at transcriptomics with space within a cell and look at where it's localizing. (00:12:05) We can look at translucent zebrafish and look at the development across different cells and when the brain develops. (00:12:13) We have sensors that allow us to look at cell-cell communication in different molecules. (00:12:18) And so we can be strategic about (00:12:22) the types of experiments and data we want to collect that helps us bridge across these, that makes it so that there's some connective tissue that helps drive the modeling that, you know, the modeling magic that happens. (00:12:35) Yeah, the reason I ask the question, by the way, is I used to be a biologist. (00:12:38) I have a PhD in biology and I worked in wet labs for almost a decade and everything else. (00:12:41) Are you looking for a job? (00:12:46) We can talk about that later. (00:12:49) It's not a no. (00:12:49) At this point in my career. (00:12:54) I'm like Danny Glover, you know, and I'm almost at retirement. (00:12:58) But I think, you know, one of the things that was always lacking was this integrative nature across the different layers of biology and the developmental biologists would work on their own, the molecular biologists would be doing different experiments. (00:13:08) And so that's what I was curious about. (00:13:10) Typically, there's a reductionist view of biology and there's a systems view. (00:13:13) And those people didn't really work together deeply. (00:13:15) And so one of the exciting things about what you're doing actually is how you're bridging that. (00:13:18) And so that was kind of the basis for the question as well. (00:13:21) Yeah, and if I could add something there, you know, it's, I think that, you know, we're in the age of this kind of information theory in biology. (00:13:27) And so (00:13:29) there are levels of complexity and hierarchy and biology and kind of each level is made-up of and, constituted by the lower levels. (00:13:39) And so as you want to have that kind of more complete description and you want to have systems that can really generalize and begin to actually answer, experimental questions digitally that you could ask in the lab, you need to have kind of the right basis for modeling at every level. (00:13:55) And so I think what's really unique about what we can do is to (00:13:58) as Priscilla and Mark were saying, really build information at each of these different layers, collect them, collect kind of those connection points, but then also really kind of do it at the scale that will reveal that underlying information architecture. (00:14:15) And that's going to be really critical to actually be able to build digital representations that can answer new experimental questions. (00:14:22) One of the things that inspires me most about this effort is really what Priscilla said, which is like, well, there's so much we actually don't understand about biology and what if we could, which I think is actually very different from lots of other incredibly interesting and useful AI problems we attack where we're like trying to replicate human behavior. (00:14:39) And I'm like, a lot of that data is, you know, on the internet or captured. (00:14:44) without pretending to understand all human behavior, you can predict a lot of it. (00:14:48) I thought one of the most interesting things in your release was actually, the like mechanistic interpretability stuff you alluded to, which is, can we actually extract new knowledge from, you know, what the model believes is happening, right? (00:15:02) Can you talk a little bit about that? (00:15:03) Yeah, I'm really excited about that. (00:15:05) So I think, in mechanistic interpretability, kind of traditionally it's been applied to large language models with the goal of understanding, kind of what is the representation space of a large language model? (00:15:16) How does it compute things? (00:15:18) And does that really connect to, you know, what we understand about our intuitive understanding of the world? (00:15:25) And so there's, I think, this really rich toolkit that has been developed to start to be able to ask those questions. (00:15:32) So kind of what does that mean for biology? (00:15:35) One of the classes of models that we train are these protein language models. (00:15:39) So they're really, you know, it's trained on the codes of proteins. (00:15:42) And so anything they learn about biology is kind of emergent. (00:15:46) And we've seen that they can learn things like biological structure and biological function. (00:15:50) And that's just kind of emergent from this, you know, token prediction training task. (00:15:55) So, as we think about mechanistic interpretability in those models, we're really seeing the unknown because the models have been trained on billions of protein sequences. (00:16:06) They've been trained on, you know, both known and unknown biology. (00:16:09) And yet they're developing these representations that start to kind of capture things that we can really see correspond to that reductive picture of biology that's been built up over the centuries. (00:16:20) So kind of you can start to connect the dots between proteins where we kind of really don't know anything about them with proteins where we do know something because there's that kind of underlying structure grammar that's linking them in the representation space of the model. (00:16:37) And at the extreme, it could be, we're going to understand systems in the body that we didn't before or the mechanism of action for a new treatment because we can ask the model, right, interrogate that representation. (00:16:48) That's right. (00:16:49) The hope is that you kind of really learn the underlying basis for how it's making the predictions. (00:16:53) And so you open up the black box and you can actually understand kind of the biology that the model is representing. (00:16:59) So asking for a friend, you know, you guys all believe in venture-backed companies as a way to have impact. (00:17:07) on the world. (00:17:09) What was it like collecting data on zebrafish or the span of the data or the wet lab work or just the scale? (00:17:17) Like what makes this a better fit for this big non-profit ecosystem effort versus a venture-backed company? (00:17:25) Well, I think we just want to give tools to the whole scientific community. (00:17:28) And I mean, (00:17:29) So I think in order to have the biggest impact, I mean, part of it is just we're, I mean, it's not actually clear that we couldn't run it as a business if we wanted to. (00:17:39) I just think that we'll have a bigger impact by getting this in more scientists' hands quicker by doing it as open source projects instead. (00:17:48) So yeah, I mean, I think that that's kind of the approach. (00:17:52) But (00:17:53) I don't know. (00:17:53) It's an interesting question. (00:17:54) I'm not sure that, I mean, obviously you were doing it as a nonprofit, as a for-profit company, a bunch of the modeling before. (00:18:02) Then you run into certain issues. (00:18:03) I mean, you have to raise a large amount of money in order to build a compute clusters. (00:18:07) You know, I mean, I think in a lot of ways the data is actually even more of a constraint. (00:18:12) And because if you look at like the scale of these models compared to language models, they're smaller, but they're smaller because the amount of data is less. (00:18:21) In order to get the data, it's not just like there's some factory somewhere that you can pay to produce the data. (00:18:29) Like you actually need to invent new novel scientific approaches to be able to do the, you know, for example, the type of cellular engineering we're doing in New York or the types of devices in Chicago, which is why, you know, when we're talking about this concept of frontier biology and frontier AI, the frontier biology is you need to do real science to advance different biological methods in order to be able to observe the things (00:18:51) that create the data that go into the model. (00:18:54) So it's not just like an off-the-shelf thing that you can create. (00:18:56) Now, that's a pretty big effort. (00:18:58) I don't know that there are like that many things like that are done as biotechs. (00:19:03) I think it's just the scale of the ambition of what we're doing, the horizon over which we're committed to doing it. (00:19:10) I think part of the theory is like, if you're building tools that are this complicated, you kind of want to have a 10 to 15 year time horizon on building out these efforts. (00:19:19) And then the scale of capital required, I mean, I guess there's no rule that said that you couldn't do it as like an incredibly well-funded startup, but I think that this just made more sense. (00:19:29) And then it also is simplifying strategically to not have to think about how you're going to make money with the different things. (00:19:36) I mean, we just, we want to get the models in people's hands. (00:19:39) We release them as open source. (00:19:41) I think that that's like a very valuable thing to do. (00:19:43) And again, I mean, the theory isn't that we're going to cure the diseases. (00:19:46) We're not. (00:19:47) It's that we want to help accelerate the pace of progress for the whole scientific field. (00:19:51) As the person least experienced with making money here, I would say that there, the sort of neutral nonprofit nature of our work actually helps harness more people to enter this effort. (00:20:05) And to actually achieve the mission of understanding the totality of human biology and to cure, prevent, manage all disease, you actually do need the entire academic biotech industry to come together and to work on this in a sort of unified way, in part because there's a lot of talent out there and it's not helpful to leave any talent, exclude any talent from the effort. (00:20:31) And there's a super long tail of diseases. (00:20:35) There are the common ones, and even the common ones, I think if you unbundle heart disease, cancer, neurodegenerative diseases, even if you unbundle like dementia or depression, there are many, many, many subcategories that become more and more niche. (00:20:52) And that's not even looking at the long, tail of rare diseases. (00:20:56) Those often get orphaned and don't get brought along when we're sort of looking at what the most efficient way to impact the lives of many. (00:21:04) But if you sort of decentralize the effort and put the tools in many people's hands, you start getting people who are like, you know what? (00:21:11) I am super interested in spinal muscular atrophy, and that's something I care deeply about. (00:21:16) And if you put the tools in that person's hands, they're going to be able to make progress. (00:21:21) In a way, if you had to focus your efforts and make big bets, (00:21:25) You probably wouldn't because it's just a niche, individual, small group disease that actually will in turn, if we can understand that disease process, helps us unlock knowledge about a lot more about how the human body works. (00:21:42) Do you have any thoughts or predictions in terms of what disease areas this work will impact first? (00:21:47) I know it's very hard to be predictive about these things, but just given the nature of the work and the nature of the models, are there areas you're most optimistic about in the short to medium term? (00:21:55) That's actually not how I think about it, at least. (00:21:58) The way I think about it is like, we want to understand how biology works. (00:22:02) The ideal world is you would say, I understand the genetics of this person. (00:22:09) So I want to think about people at the individual level. (00:22:11) I want to understand the genetics of this person. (00:22:13) I want to understand the risks they have to different illnesses. (00:22:17) I want to understand the mechanistic connection between, say, (00:22:23) a gene variant, a protein, and a disease process. (00:22:27) Because if you understand that through chain, then you can design a protein, design a drug, bespoke to them, and actually make an intervention. (00:22:35) And right now, I'm sure we've all had experiences being sick. (00:22:40) And if you have something that's even remotely non-standard, you go into PubMed, you look up a paper, you look up the supplement, and then you start going through the methods, and you're like, am I represented in this paper? (00:22:54) And we're just making guesses. (00:22:56) We really have no mechanistic understanding. (00:22:59) We're saying like, okay, you're kind of like these people that we studied, and this drug kind of impacts the pathway that we think is implicated. (00:23:09) Let's try and see if anything happens. (00:23:12) And time passes, and sometimes it works and sometimes it doesn't. (00:23:17) So my goal is to be able to treat the individual as an individual, understand the mechanisms, and be able to intervene. (00:23:25) And there are different diseases that are at different stages of filling out that whole through line. (00:23:31) And so for some diseases, you just want to understand which gene variants actually cause disease and which don't. (00:23:40) And that in itself can be super empowering to patients. (00:23:44) And if beyond that, there are some diseases where we understand the chain, we just can't intervene and change a specific protein function. (00:23:55) That's super exciting too. (00:23:56) Like if we could design a protein to

Around this claim