If you can completely sort of really understand the workings of a cell, that will have amazing consequences for the fight against diseases like cancer. >> I feel like I've been dropped into a science fiction movie. Like we we could only dream of of having technology like this when I was a child. Now I also have a question and I was thinking that maybe only you can answer this question. >> What Alpha Genome Atlas does, it basically premputes all of this. Alpha genome atlas is 30 times the size of alpha fold database. Wow. So the scale at which uh alpha genome can produce this the the detailed information about the effect of [clears throat] variation is unprecedented. >> Good to see you push. Look this is really challenging for me because this is very much outside my field of expertise >> if I have any at all. As always I come here today as a student to learn from you. What is alpha genome? >> Alpha genome is a machine learning model which makes predictions about variation in the genome. So our genome uh not just the human genome but uh genome is basically the recipe of life. Uh what it is it's essentially a code. In the case of humans, it's a code which has three billion characters. And if one of those characters changes, what would be the effect on the actual cell properties? What will be the effect on the organism? These are the things that alpha genome is trying to sort of uncover. >> And if you're the person on the street, what could alpha genome change eventually about medicine or their life? Everyone sort of is affected by uh what happens at the genomics level. Uh the genome is the recipe of life. It has both the ingredient section which says which proteins uh are expressed and the sort of mixing section which sort of talks about uh gene regulation uh when do these proteins get expressed where do they get expressed and in what quantity >> and all of that information is encoded in this recipe of life. Every single person on the planet has a recipe. I have a recipe. You have a recipe. The person on the street has a recipe. All of our recipes, while they are similar, there are some changes. But what do these changes actually mean? What are the implications of these changes? That's what we have been trying to sort of understand for many many years if not decades. With the human genome project for the first time we were able to read the entire recipe book at least for one human sort of genome. U but now we are trying to figure out how do we decipher it or understand what actually happens if you make a change in one particular character. What are the consequences of it? Will it cause a specific disease? Will it basically be uh connected to uh high susceptibility for some kind of uh u problematic issue like maybe having a high body mass index or having more susceptibility to certain cancers. All of these things are can be decoded in some sense by better understanding of uh the variations that happen in the genome. >> Mhm. That is incredible. And I really wanted to ask this because my problem every now and then I try to read medical papers. It's not my field of expertise but I I'm a student. I just try to learn. >> Yeah. >> And my problem is that whenever I read about I don't know any kind of medicine. My problem is that every person is different. >> Yeah. >> And I go to the barber and the barber says that he never saw two heads that are the same. Yeah. >> And I read a paper on muscle hypertrophy like like bodybuilding paper how to build muscle and they have a bunch of people carefully selected, same age, >> same height, same build, >> they get the same diet, they do the same training and one guy gets incredibly muscular. They do the measurements. >> Yeah. >> Some others get a bit of muscle. One poor guy didn't get any muscle at all. And there was one who I think lost muscle. Like it's crazy how how different we are. I'm thinking that that is it possible that that we could finally be one step closer to understanding how how everyone is a bit different and what that means. >> Yeah, exactly. I think so this understanding or deciphering this code of life, what are the consequences? Sometimes those consequences might not mean anything, right? Like sometimes the changes are like uh they are not that material. Sometimes they are very material sort of consequences which are very well understood. Right? For example, cickle cell anemia. There is one particular sort of mutation which causes sort of cickle cell anemia and we understand what happens there. But most of the things that sort of you care about these days u requ are sort of polygenic in the sense that there are multiple sort of uh mutations that are uh acting on that are causing that particular uh incident or that that particular phenomena. And in some sense what alpha genome tries to do is tries to figure out like what happens at especially at a single sort of uh uh mutation variation level uh at the molecular sort of phenotypes. Now how do those molecular phenotypes uh then uh get affect the the property of the whole organism or in in the case of humans the the what the human sort of feels like the s susceptibility to specific disease that's something that will still need to be sort of connected and uh and research will need to be sort of done on that. >> My understanding is that before we dive deeper alpha genome is free for all of us forever for non-commercial use. Yes, exactly. So for as we have sort of done uh for many of our uh scientific uh models whether it's alpha fold one alpha fold 2 u alpha genome we have made it accessible to the entire scientific community for academic sort of use uh for them to sort of use it for their own sort of research fine-tune it sort of and the model weights are available. Are these true? And please explain. >> Yeah. >> The genome is the instruction manual for cells. >> Absolutely. I think the uh the genome as I was saying it's like a recipe book. Uh it has both the coding part of the genome which is like the ingredients sort of section of the recipe book book which says uh what are the proteins that will need to be synthesized. So the human proteome has 20,000 proteins. So our our genome sort of encodes these 20,000 sort of proteins and there could be variations in those 20,000 proteins. And then there's the non-coding part which is the which is sort of which looks at how are these proteins uh sort of mixed. How are these proteins sort of expressed? where are they expressed in what quantity and that part is basically the dark part of the genome which is much less uh well understood because it sort of talks about the regulation of uh uh of genes and uh protein expression um and so I think what alpha genome is trying to do is trying to sort of uh uh sort of decipher and inform this this whole sort of space. Now I also have a question and I was thinking that maybe only you can answer this question. So I'm I'm really curious about this. This is trained on human and mouse genomes. >> Yes. >> And what I would love to know is that you train it on the human genome. It has some amount of knowledge and and and quality of of results. Then you train it on the mouse genome too. after that does it get better at the human part? >> Yeah. And so this I think currently I mean you have sort of sort of seen this in many different sort of places in machine learning neural networks try to fit to the data and uh if you just give it the human data uh the neural network has this susceptibility that it can memorize or sort of uh uh think about uh the mechanisms that only work in the case of the human genome. But by training it by forcing it forcing the model to explain the same effects in the context of the the mouse genome, we are asking the the model to think broader and think about general sort of principles that can explain both these different types of variations. And that really helps in not only model generalization but also making the model sort of learn about uh sort of what really matters and what's really happening at the physiological level uh sort of in those molecular phenotypes. I've seen it in another Google deep mind paper that was the SEMA project which played a bunch of video games and I really found it stunning that you play a video game you have some score on it and after playing a bunch of other video games you get better at this game >> I think >> and it seems like maybe the same phenomenon that that but that sounds like intelligence to me right you you you apply the knowledge from other things to to this thing >> yeah and I think that's transfer learning in action where you are able to extract concepts from things which are related but are different problems and it helps you become better. Right? So if you sort of uh solve your maths problems and then solve your physics problems but just getting better on maths helps you get better on physics. My understanding is that the big advance here was that previous models were either too zoomed in and shortsighted or they looked from far away but at a low resolution. Now this one can do both. >> Yes. >> True or false? >> This is true. So when we started on this journey of u going after this extremely ambitious challenge of deciphering the genome um there were a lot of technical challenges. uh our genome as I just mentioned is 3 billion sort of uh base pairs ATCG right uh sort of a very long sort of sequence and if you want to understand what is the effect of a particular sort of variation you need to look at a broad uh sort of an expanded uh neighborhood around that uh that variant because it can affect sort of uh uh things which are quite far off because of 3D interactions and so on. It's not that a particular sort of base pair only affects uh what happens in its neighborhood like it the neighborhood could be very large. So earlier sort of variations of the model. So we had a model called informer which preceded alpha genome where both the context window in which we were looking at uh the neighborhood was sort of smaller as well as the resolution at which we were trying to analyze the genome was much coarser. Right? We were not looking at what's happening at the base pair level. We were sort of looking at sort of a large number of sort of uh base pairs taken together and we were not looking at 1 million base pair window. We were looking at a smaller sort of window. What alpha genome I tries to do and is the first model to achieve that is to operate at the base pair resolution and have this very large window and uh and that allows it to not only make sort of prediction at a lot of precision which require a lot of precision precision but also uh sort of make predictions which involve reasoning about a much larger set of uh uh things that are happening in the neighborhood of that uh variant. I saw that it is so dominant it matched or beat the strongest comparison model in 25 out of 26 variant effect benchmarks. I think only at alpha fold did I hear like this kind of dominance that is absolutely incredible. Now what I'd like to ask you when you saw that it first started working >> Mhm. >> what did it feel like? >> Yeah. So I think one thing that um we were we have been working on this problem for a long time as with many of our scientific uh challenges. Um it is these are incredibly hard problems um and require innovation in many different things like data and sort of uh coming up with a better architecture coming up with sort of new ideas on training and so forth. When we started seeing these results, we knew uh that we were in a different phase and uh part of it was expected because we had done all the hard work of uh enabling training to happen at that resolution which had never been sort of done before. But to get it right and to actually show that uh the original hypothesis that we had gone in with that if you increase the resolution, if you increase the context window, if you regularize the training properly, then neural networks will sort of uh show pro uh show the right results and so that was really rewarding uh to see for the whole team. Was there a phase of suspicion like it it it can't be right you know maybe maybe we we tested on the training set or or something. >> Yeah. So I think this is where we have uh our team sort of we are very very cautious scientists. So uh there is a lot of work that the team does on sort of uh on metrics and evaluation and so when the results came in there's you so you there's a genuine sort of um uh sort of expectation that you you're going to uh uh check it again to make sure that everything is working properly and these are sort of valid results but uh we had a lot of confidence in our metrics that we had set them set them correctly because it was not something that we we have been doing for the last sort of 3 months like it's it has been a project for the last almost four five years or more. >> What is alpha genome atlas? >> So alpha genome atlas is essentially uh a sort of a dictionary of what happens when you have variation in the genome. So now I told you about alpha genome. What alpha genome does is it looks at um your genome and you can ask it the question if I change this uh base pair or this character in your genome from a to a t or t to a c what would be the consequence and that's very helpful but suppose you wanted to understand the effects of all variations now all variation there are a lot of variations because like there are three billion base pairs and all of them can be changed three sort of times, right? A A could be changed into a T or a C or a G, right? So there are 9 billion possible single nucleotide variations that you can study. And for every single one of them, there's a lot of data that you can sort of uh that you can extract as to what will be its effect at the molecular level at uh the effect on splicing, the effect on sort of gene expression and and so forth. What alpha genome atlas does, it basically precomputes all of this. >> It basically looks at what happens for all these variations and makes this massive data sets a pabyte size sort of data set and makes it available to the scientific community. >> That's unbelievable. It feels like this alpha fold moment where the inference of time of folding became you know so quick. Yeah. That at some point Damis just said that you know let's fold all the proteins like atlas sounds like that for the genome. >> Exactly. So I mean for for alpha alpha fold we had um we made the predictions for uh structures of 250 million proteins almost all the proteins that were known at that sort of time. But even the alpha fold database which has now been used uh by more than 4 million users across 190 countries alpha in terms of its actual size alpha genome atlas is 30 times the size of alpha fold database. >> Wow. >> So the scale at which uh alpha genome can produce this the the detailed information about the effect of variation is unprecedented. All this from one megabase of DNA sequence. >> Yeah. So with with a 1 million sort of uh base pair window uh in in context >> for mutations that affect gene activity there is a fairly strong correlation between the predicted and the observed effect size. I saw in the paper that row is about 0.5. >> Mhm. >> But but that's stunning. >> Mhm. >> That's that's incredible. And uh my understanding is also that this doesn't just predict whether a mutation matters but often gets the magnitude of the effect roughly correct too. Is that fair to say? >> Yes, absolutely. I think what we did with the especially alpha genome atlas is we came up with um a new way to sort of help scientists sort of look at the data and explore this mass amount of data. So the alpha genome variant impact score or the AVI score is uh sort of carefully sort of calibrated to say if the score is high that means that there is that particular variant is a variant that you should really think about and analyze uh sort of carefully. Um and if the AVS code is low, you can there's high confidence that it will not have uh a dramatic impact on various different sort of phenotypes. >> Mhm. So to me, this sounds like one of those root node problems where you you solve this problem and then it's the applications can grow out of it in so many directions. I check alpha fold every now and then which was one of these root node problems and in the first couple years >> I could actually look through the papers that that used it and I tried to understand what they do now there is so much that's built on top of alpha fold that it's completely hopeless to for any of us to understand all of it does it look like alpha genome is going to be the same >> yeah I think alpha genome with alpha genome we are trying to solve that root node problem of deciphering the genome, right? And and as you sort of mentioned, it it is really truly a root node problems node problem in terms of its implications. So not only does it have implications for understanding sort of uh diseases better, it also has implications for how would you treat diseases. Uh what are the sort of key things that you need to sort of understand about biology and what are the places where you can sort of control expression or what what variations sort of control expression uh of certain sort of proteins which are relevant in that disease path pathway. But it's not just uh applicable for um understanding uh diseases and human health. Uh it has implications that go beyond in uh in broader sort of biology. So in fact uh sort of now we are seeing um a lot of interest in use of generative AI in synthetic biology and again when you are sort of synthesizing sort of uh u new organisms or new sort of proteins like what are the properties that uh we should be sort of thinking of right like you would have seen uh DNA language models like EVO uh sort of come up and and people using them to sample different genomes to uh to solve different types of problems and again alpha genome gives you sort of a lens by which you can study the properties of those genomes and so uh the implications of uh this model and its future sort of variations which we are working on I think will be sort of uh in the same way um very broad >> I feel like I've been dropped into a science science fiction movie like we we could only dream of of having technology like this when I was a child that's I think that's unbelievable how do we know that it learned mechanisms not just correlations like usually that's the that's the problem with many of these papers that the difficult part is actually proving the connection between the two that's almost impossible in many cases >> yeah so I think that's a very good question I think and that has been one of the sort of challenges right when you are looking in computational biology that you don't want just to learn all the statistical correlations. You want to have uh the model you want the model to be able to sort of understand the actual sort of causal sort of links that happen in uh that make the predictions possible. And uh if you think about it like in alpha genome how that is sort of encoded one is that is in terms of the properties that it predicts. It predicts properties at the molecular sort of level. So it can you can really see when it makes a prediction about a particular variation about say splice sites or splicing or gene expression. So those are very sort of uh uh uh those are predictions which are which can be explained biohysically right rather than sort of corrationally that oh this variation happened basically causes that particular sort of disease but this alpha genome is not saying that it's basically actually telling you what is happening at the cellular sort of level and then in terms of validating uh what we have seen is actually when you make take these predictions and actually try to validate them in the lab. We see that those predictions make sense and you can make novel sort of findings uh of uh of variation and and importance of certain variants >> essentially because the whole thing is based in physics. This is what essentially helps in terms of verification. >> Yeah, I think the tasks are more physically required to reason about the physics. Now how the model is doing it that's uh something that we we are still sort of uh looking at and interpreting but uh uh the the the tasks that it is being asked to solve are grounded in physics. Now you have this incredibly sophisticated learning algorithm in alpha genome and then in the paper on top of it you have this tiny super simple layer of lasso regularization and I found that really amusing. I imagine this almost like a a Formula 1 engine you invented and it has a a little bicycle bell on top. >> Yeah, I I think that this sort of shows the the power of the representation, right? what you want uh models like alpha genome to exhibit is sort of a a fundamental understanding of what they are um what they are seeing or how they are perceiving uh the DNA sequence. So in in some sense most of the work should be done by the model in terms of the feature representation and then the sort of the final sort of predictor on top should take a very simple yeah almost trivial form because most of the work has happened at the representation sort of uh layer. >> What's the next step? So there are so many things that uh uh we are working on which is like with alpha genome the the story does not end here like there is a lot that we need to do to improve alpha genome to make it even uh better uh in terms of its understanding of um the genome. Um if you look at genomics and uh biology, biology sort of has a lot of different contexts. For example, different cells behave differently, right? So uh making sort of alpha genome more uh sort of aware of the context in which it needs to make these predictions is one thing that we are sort of trying to work on. um expanding the amount of data that alpha genome can be trained on so that it can uh sort of become much more accurate uh about different kinds of properties connecting these molecular phenotypes that alpha genome uh uh is able to sort of produce and uh showing and uh showing their connections and impact on properties at the organism sort of at the human sort of level right the propensity to disease and so on. How do we sort of facilitate that work? All these things are sort of directions that the team is now looking at. >> There's so much going on. I I don't want to get greedy, but I just wanted to hear about the next steps like what what are you building on on top of this? That's that's awesome. You use Gemini with your son to explore science questions like how airplanes generate thrust. [gasps] Has Gemini ever explained something to your son better than you could? So where you said, "Oh man, that's that's way better than what I would have said." >> Yeah. [laughter] So I I think it's it's not u only the ability of these models in explaining concepts but by using different analogies and so on but also uh the ability of these models to actually uh show the theory in action. So I was um a few sort of weeks back uh I I was trying to sort of explain to my son like what happens in a wind tunnel when sort of uh engineers are designing uh sort of u say um airplane uh sort of parts and what are the effects and how do they do these simulations and uh we we asked Gemini to explain it and it it in fact it creat created a simulation and where sort of uh you could actually play around with different parameters and see what how the airflow sort of uh changes. Now it was not completely accurate but uh but it was a very good demonstration of what is possible right uh and like and creating these visual artifacts that help uh someone understand the concept. We are living in an incredible time because I I also remember when I first wrote my Fav Stokes physics simulation for for fluids. I remember when I first looked at the theory and I thought I'm never going to understand this. This is this is never happening. And it took months and months of studying and and trying and now I I understand it very little about it. But at least a basic simulator I can I can write. And you just here you just enter a prompt and you just got a simulation by the way on the side. >> Yeah. >> Because the because the intelligence is so incredible that it it can do this sort of thing like like just make up a wind tunnel test and and you can probably running real time you can see the streamlines as as it's unbelievable and it's only going to get better. >> Yeah. Now I think we had about a 1 millionx AI speed up in the last 10 years. If you factor in everything, it's not just hardware. It's if you factor in, I don't know, transformers, >> how much better the implementations got, the engineering, the everything. Let's say we got a 1 millionx in the last 10 years. >> If in the next 10 years we have another 1 millionx, >> what could we do? There are so many uh open scientific challenges that are beyond our current capabilities. Simulating even a cell, creating a virtual cell, imagine the consequences of that, right? Like we would if you can completely sort of really understand the workings of a cell that will have amazing consequences for the fight against diseases like cancer or synthetic biology and so forth. Now imagine not and that's we are talking about a cell right imagine sort of simulating a organism that will be sort of that is the power that can come from that understanding and that capability but there's so many things um about the world that we still are scratching the surface on sort of understanding uh long-term climate for instance uh what are going to be the consequences of certain um sort of actions on long-term sort of climate. Um sort of coming up with new material types that can help us sort of store energy sort of uh transform energy in different forms in a much more efficient manner. understanding biology at uh from first principles to being able to truly design novel function for tackling uh some of the challenges that we are seeing both in human health as well as sort of in agriculture and many other uh things that the world is facing. There are so many sort of um problems that can be unlocked with that progress. >> That was fantastic. Usually I ask this question people and I see that it kind of breaks their brain too because 1 millionx like 10 years ago like 2016 we had no idea what what we could do today and and it's it's so exciting to think that okay if you had another 10 years what would we be capable of and everyone is thinking so small but but this one this one was really good >> nature is extremely sophisticated right sort of we are and and really I I think there's a lot to look forward to in terms of um understanding the world. I mean understanding we are talking about understanding cells understanding our own brain like uh is another sort of challenge. What is actually happening um in uh sort of materials? Why are they exhibiting certain properties? We even sort of properties that we uh have seen uh forget about new types of properties but properties that we have already seen like superconductivity we don't have a complete theory about superc conductivity right we have one sort of narrow theory of superconductivity imagine the new properties that we are yet to sort of explore or uh sort of really uh crisply understand why these phenomenas happen and the ability to be able to design those objects. Fantastic. Push meat. Every time we meet, I learned so much. And you know what the best part is? I I will go home and I cut these videos together myself. So, I'm going to watch it like five or six more times and I guarantee you that every lesson there's going to be something new that I learned. So, this was amazing. Thank you so much. >> Thank you. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one. Run inference or text to image or video. Easy peasy. Running a Deepseek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and [clears throat] moments later, results. Love it. Seriously, try it out now at lambda.ai/papers. AI/ Peepers
Generated algorithmically for Search Engine
Indexing.