While we have a tendency to simply think of protein as food, the complex biomolecules that are buried within steak and chicken are far more interesting than that. In order to better understand these organisms, scientists have recently created an advanced 3D modeling software that takes a simple text input and generates a fully interactable 3D protein. This software is called Alphafold. However, before we explore the incredible breakthroughs that have been made in protein modeling, we should examine exactly what they are beforehand.
Proteins are large, complex biological molecules that act as essential nano-machines, facilitating cellular processes by interacting with and manipulating the environment around them in response to stimuli.
Essentially, think of proteins as workers on an assembly line and the end result is the product. Proteins are also designed to have a precise, systematic order. All are created as a long, helix-shaped string formed from the combination of a carbon atom, a hydrogen atom, an amine group, a carboxyl group (two smaller molecules), and finally one of twenty different side chains. Each one of these side chains determines what amino acid the molecule is. By combining the individual amino acids in structured bonds called polypeptides, as well as twisting and bending these chains in atomically precise ways, we end up with a functional protein. The most important thing to understand here is how the shape of the protein affects it. These chains live and die by a system in which their helices and beta sheets interlock in each other’s grooves. However, if the proteins don’t work correctly, any given bodily process could be impeded or even made impossible, depending on the severity of the defect.
This can have serious, life threatening consequences along with being incredibly hard to treat, as the disease is not an infection, but a mistake in the person’s genetic code. Some diseases you may have heard of that fall under this are Alzheimer’s, Parkinson’s, Type 2 Diabetes and Cystic Fibrosis. Given this fact, it has always been an overarching goal of the biological sciences to understand how proteins are structured. However, the field has also been constantly plagued by realities which seemed to stop us from ever achieving that goal.
Given the vital importance of proteins to the functioning of the body, it’s only natural that scientists were incredibly curious about what the structure of the protein could look like. In 1951, they made a breakthrough. The trio of Linus Pauling, Robert Corey, and Herman Branson was the first to document alpha and beta sheets, the backbone of protein formation. However, it would be a few more years before the first full structure would be solved. This would come in 1957, when John Kendrew and Max Perutz discovered a method, known as crystallography, to reliably anticipate protein structures by soaking crystallized proteins in heavy metal atoms, creating a dependable diffraction pattern that could be used to predict physical structure.
From this point onward, the craft of crystallography would develop substantially for the next three to four decades, condensing a nearly 20 year long process to determine a single protein to one that would take only three to five years by the mid 1990s. The biggest problem with this methodology was the sheer scale of potential proteins and the limits of the time. As said earlier, the shape of a protein is the most important part of their function. A simple peptide bond with 35 molecules on its chain can have up to 3.4 x 10^45 possible folding combinations, each with a different purpose, though most with none at all, in the body. However, using our understanding of chemical bonds, how certain types of folds have higher probability of binding given XYZ conditions, and which bonds and folds would violate energy stability and cause thermodynamic disruption, ideas were floated about potentially using computer learning to try to crack the case. This idea had many detractors, with people using mathematics to argue that even given millions of years, linear machine computation could never definitively solve the protein problem. Still, it caused enough of a stir to start something.
It is in 1994 that we saw the first new approach to protein prediction since its inception in the 50s. This kicked off with the first Critical Assessment of protein Structure Prediction, or CASP. Organized by the team of John Moult, Krzysztof Fidelis, Jan Pedersen, and Richard Judson, the goal of the project was for groups of computer scientists and researchers to compete at creating software that would accurately predict protein structures. Contestants would have a year beforehand to create a computer modeling system which would be able to accurately predict the shape of a protein simply based on its amino acid code. The output prediction would then be referenced against the crystallography solved structure and graded out of 100 for accuracy. While 100 was the goal, the designers of the program knew that was completely unfeasible, so they set the target at 90. CASP would end up innovating in the field of protein modeling and prediction in many interesting ways, with one standout being the game ‘Foldit,’ a game in which players try to assemble the most accurate protein model possible and pool their results to see who could be the most accurate. However, the greatest innovation at CASP was yet to come.
At the 2018 CASP summit, we saw the first usage of artificial intelligence and deep learning techniques for protein folding. This was Alphafold One (AF1), a novel protein prediction software developed by the DeepMind team, lead by Demis Hassabis. Alphafold was developed during the 2010s, as AI deep learning techniques and their applications in geometric predictions increased in scale and capacity. As anyone familiar with artificial intelligence knows, the biggest bottleneck in a deep learning system is a paucity of data, which is a problem the DeepMind team luckily did not have to deal with due to the work of the Protein Data Bank (PDB). The PDB is a globally connected database of all known proteins and as many of their associated structures as has been verified by labs.
Using the data from the PDB, they trained models to look at a simple 2D nucleotide sequence and from there, infer information about how the DNA would be translated into geometry and function to the point where a fully rendered 3D model of the protein could be made from a 2D sequence alone, completely by an autonomous machine. To give the AI a medium to undertake these calculations, (and as the initial inspiration for the whole project) Hassabis looked back at the game Foldit and wanted to see if he and his team could train a deep learning system to play the game even better than humans could. To increase the model’s efficacy, the team also created a separate algorithm that was trained specifically on amino acid grouping dynamics and the ways they attract/bond. Like nucleotides under them, amino acids also have a preferential connective process, similar to nucleotides. Positive side chains attract negative, hydrophobic and hydrophilic (resistant/absorbant to water) attract like molecules. By training a system to correlate those pairs and understand the fundamental physical dynamics which make a peptide bond stable, they eliminated a lot of the mistakes that an AI could make playing Foldit by not properly accounting for those dynamics and relationships. Once this data was fed into a 3D rendering model, the DeepMind team would let this larger model bend and fold the peptide based on attraction, distance and realistic torsion. With their approach, DeepMind saw tremendous success, achieving the highest CASP score in recent competition history at that time with a 70. Despite this achievement, not hitting the 90 percentile target stung for the team, which only meant that more progress needed to be made.
After the supposed underperformance of Alphafold 1, the DeepMind team decided that for the second iteration, the entire structure of the system needed to be different. Instead of having an external AI to analyze geometric data and probable folding patterns, the team wanted to integrate that system directly into the 3D modeling software which they hypothesized would increase performance exponentially. In AF1, the software would process the probable links between amino acids and then send that information to be optimized into a 3D structure by a different model. The potential for error here is part of what held AF1 back. Alphafold 2(AF2), by comparison, runs this process both completely internally and simultaneously, meaning that the room for error was reduced, as the model’s internal logic was one consistent chain, not two models communicating two separate processing tracks. This type of novel analysis was done by a machine called the Evoformer, a machine developed completely by the DeepMind team.
The Evoformer is split into two ‘towers’ or datasets. The first one focused on the biological structure of the protein based on its amino acid sequence. The second one, the geometric tower, used protein structure data to predict how those amino acids would fold sequentially. By having the two towers build on each other through constant iteration, the Deepmind team found the most accurate means of protein prediction to that point. For example, if the geometry tower finds that a link the bio tower sent over is physically impossible, it tells the bio tower to ignore it. Similarly, if the geometry tower predicts two positively charged side chains attracting to one another, the bio tower will correct that prediction in the final output. By having a stronger logic chain to correct errors and catch mistakes, the DeepMind team effectively made an AI that could be critical of its own work in real time as it’s completing a task. After the final input is collected between the two towers, the 3D rendering model takes it and renders it and then the process is repeated three to four times to create slightly different versions of the protein and the most aggregate of those outputs will be the final output.
During the CASP 14 summit in December of 2020, DeepMind finally did it. Alphafold 2 entered the competition as a question mark and exited with a performance score of 92.4, shooting past the benchmarks set by any other protein prediction software. This was the first time anyone had broken the benchmark and was an absolute jawdropper for the bio-science community. To put how revolutionary Alphafold was into context: since its inception in the 50s, crystallographers in labs around the world had painstakingly worked to determine the structure of 150,000 proteins in 2020. Alphafold was able to synthesize over 200 million. For the progress that the team leaders were able to make in solving this seemingly unsolvable problem, Demis Hassabis and John Jumper were awarded the Nobel Prize in Chemistry for their work on Alphafold 2 in 2024.
Alphafold’s practical medical usage is unparalleled. By giving us a window into the precise structure of untold proteins, scientists have been able to create a vaccine for malaria by identifying novel structural elements. Alphafold is also used in order to combat antibiotic resistant bacteria, which can be fatal as it won’t respond to conventional antibiotic treatment. The technology has been used for research into Parkinson’s, cancer, and so many more diseases that will create positive impacts for society. In one specific example, labs were able to design proteins that emulated antibodies to diseases cultivated in non-human organisms, which helps to eliminate the risk of allergic reactions patients can have from having non-human antibodies injected in them by tricking the body into ignoring the fact the antibodies have a non-native origin.
Finally, it’s not as if the DeepMind team has stopped now that they seemingly solved the problem of protein folding. Armed with the compute power of Google and scientific minds across the world, the team continued iterating on the idea of Evoformer technology, eventually scrapping it for a new approach all together. This would become Alphafold 3, the current version of Alphafold. The problem behind AF3 was that while AF2 was incredibly good at predicting individual protein structure and multiple protein structure, due to the specific setup of the Evoformer, the system could not process any non-amino acid based interaction in the protein complex. This means that essential interactions with molecules like RNA and ligands (the scientific name for any molecule that binds to a protein, such as ATP or dopamine) could not be modeled. This limited the real world applicability of Alphafold severely, as for many essential proteins, understanding how the binding to non-biological molecules works is the key element in the system functioning.
In order to deal with this, the Evoformer was scrapped and a new system was built from the ground up. Instead of building proteins from scratch, the team would give an AI model a solved protein structure that had significant amounts of random statistical noise in it. The AI’s job was then to eliminate the noise so the user is left with an accurate and complete protein. This system doesn’t rely on the biological information setup from the Evoformer, which is what allows non-biological molecules to be modeled. However, the biggest problem here is the fact that most protein structures don’t have their exact structure in the PDB. In order to combat this, Alphafold uses multiple sequence alignment, which is the process of taking a protein from one species and trying to find the closest possible homolog in a different species which does have a solved structure. This works across phylogeny because the basic structures of most proteins stay remarkably consistent in our cells across millions if not billions of years of evolution. Alphafold 3 is open source and accessible to anyone, which has created opportunities for millions of people to engage with revolutionary science in the most personal way possible.
Even though Alphafold and its subsequent siblings have been a massive success in tackling some of the biggest problems in the bioscience world, the DeepMind team isn’t just limiting themselves to the protein folding problem.
In order to understand their next technological advancement, Alphagenome, we have to get a handle on how proteins are made. Proteins are not a simple molecule, despite their biological popularity. They do not simply appear and start working, but rather are built from the ground up by your DNA. DNA (deoxyribonucleaic acid) is THE fundamental object that makes us, us. It acts essentially as the instructions in a LEGO set, telling the body what to make and where to put it.
It works like this. Information is held in your DNA, which is composed of ‘base pairs,’ a set of nitrogenous bases held together by hydrogen bonds. The pairs correspond, with Adenine-Thymine and Cytosine-Guanine always pairing up. Then, these pairs are arranged in a specific sequence that conveys information, in a similar vein to written language. The smallest alteration in the intended layout could do absolutely nothing or it could permanently impair the person’s entire nervous system, depending on where it occurs in the sequence. From this inert information stage, we move onto “transcription,” the process by which DNA is turned into proteins. For this, the cell uses RNA, or ribonucleic acid, which acts essentially as a mailman for the DNA. In contrast to DNA, RNA is a much more active and unstable molecule that also disintegrates after its function is filled. The key difference between the two is in their structure, as RNA is (usually) single stranded and replaces Thymine with Uracil. This fundamental difference in the bonding structure, as well as RNA’s single stranded form, allow it to bond with DNA in order to obtain information and ‘read’ it. To ‘read’ it, mRNA, a messenger that can take DNA codes and transport them, must first open the two stranded DNA with a protein called RNA polymerase, which has the ability to unravel DNA. The unraveled strain of DNA base pairs matches up with the corresponding pairs in the mRNA, successfully recording the genetic information in the messenger for transcription. After the messenger has received its mail, the base pair bonds of the DNA will reconfigure back into their natural position for future transcription.
From here, the RNA, or more specifically the messenger RNA (mRNA), takes this information and travels to a section of the cell called ribosomes. The ribosome is essentially a protein factory, getting data inputs from mRNA and turning them into functional proteins for bodily usage. This is done in a process called translation. It starts when the mRNA docks to a small ribosomal subunit which recognizes specific mRNA signals. This ensures the correct messages dock in the correct factories. After all, a toaster factory simply won’t have the parts to build a TV. From there, the ribosome scans the mRNA one codon at a time, each of which equates to a specific amino acid. A codon is a three letter string of base nucleotides that corresponds to one of the twenty standard amino acids used in biological construction. This is a purely mathematical construction, as if codons were two letters long, they would only have 16 possible combinations while with three they can have up to 64 combinations. DNA codes are degenerate, which simply means different codons will correspond to the same amino acid. After that, transfer RNA (tRNA) molecules connected to the ribosome begin to read the DNA information off of mRNA with anticodons, oppositely attracted sets of codons which ensure correct binding. Once the tRNA hits its determined end, also known as a stop codon, it will know that the protein has been assembled and its work is done. The mRNA dissolves into the cytoplasm of the cell while helper proteins called chaperones bend the newly formed protein strands into their functional forms. From here, the protein can be used for whatever function is necessary to keep its host body alive and functioning. In this way, it is useful to think of proteins like tiny, microscopic machines. By virtue of their chemical structure and physical folding patterns, they are each assigned a specific role in the operation of bodily functions. The same way that a certain gear needs to turn in order for a conveyor belt to spin, the body needs a certain protein, in a certain place, with a certain shape to fulfill its job in order for the body to function.
AlphaGenome is tackling a completely different problem than AlphaFold. Where AlphaFold takes a protein’s amino acid sequence and predicts the 3D shape it will fold into, AlphaGenome is working a step back from that. It takes raw DNA, up to one million base pairs at a time, and tries to predict how the cell is going to regulate that stretch of genetic code. It isn’t modeling protein structure at all.
What it’s modeling is everything that happens in between having DNA and actually getting a working protein out of it: which genes get switched on or off, how much of a given gene is being expressed, whether the mRNA transcripts get spliced correctly, and in what tissues all of this occurs. To do this, the system works in two stages.
First, a set of convolutional layers, the same type of pattern recognition system used in image recognition software, scans across the DNA sequence looking for short, recurring patterns. These would be things like transcription factor binding sites or splice site signals, the small sequences that the cell’s own machinery uses to know where to act. Think of this like the model learning the vocabulary of the genome before it tries to read the sentence.
From there, that information gets passed into a transformer, the same basic type of architecture that powers large language models. The transformer is what allows the system to look at the full million base pair stretch all at once, which matters because gene regulation isn’t always local. For example, an enhancer, which is a small region of DNA that boosts the expression of a gene, can sit hundreds of thousands of base pairs away from the gene it controls and still affect it through the physical looping of the DNA strand. If your model can only see a small window, it will never catch that relationship. Older models like Enformer, AlphaGenome’s direct predecessor, were stuck in a tradeoff: they could either look at long stretches of DNA but only make rough predictions, or make precise predictions but only on short stretches. AlphaGenome broke through that wall. It can analyze a full million base pairs and still make predictions down to the resolution of a single base pair, and it did this without blowing up the computational cost.
Training a single AlphaGenome model took only four hours, which is actually half the compute budget that Enformer needed. What comes out the other end is not one prediction but thousands, run simultaneously across different molecular readouts: gene expression levels in different tissues and cell types, chromatin accessibility, histone modifications, transcription factor binding, splice junctions, and 3D chromatin contact maps. To figure out what a specific mutation does, the model runs twice. Once on the normal, healthy reference sequence and once on the mutated version. The difference between those two outputs is effectively a complete map of how that single letter changes ripples across the entire regulatory system. That is what makes AlphaGenome medically powerful. It is not diagnosing diseases we already know about. It is turning a variant that was statistically associated with a condition into a real, mechanistic explanation of what is actually going wrong, at what stage in the DNA to protein pipeline, and in which specific tissues.
This process, the whole chain from DNA to RNA to protein, didn’t just happen once when you were developing in the womb. It’s happening right now. Your cells are reading DNA, transcribing it into mRNA, splicing that mRNA and translating it into proteins as you sit here reading this. They will keep doing that until you die. This means genetic conditions caused by mutations aren’t the product of something that just went wrong a long time ago. The mutation is in your DNA right now, being read and followed by your cells on repeat, producing defective output every time. That’s actually good news. If the cell is still actively executing bad instructions, you can step into the middle of that process and correct things. You won’t reverse damage already done, but you can stop it from continuing.
The hard part has always been figuring out what exactly is going wrong and where. Sequencing a patient’s genome turns up millions of spots where their DNA differs from the reference. Almost all of those are harmless natural variations. Somewhere in that ocean, one variant is the actual cause of the condition, and figuring out what it’s doing at the molecular level has historically been brutal, case-by-case lab work. Is it knocking out a transcription factor binding site so RNA polymerase can’t find the gene? Messing up a splice site so the spliceosome produces a garbage transcript? Weakening an enhancer hundreds of thousands of base pairs away, quietly reducing expression in one tissue? Each of those is a different problem at a different step in the pipeline and each calls for a different fix. This is the exact problem AlphaGenome was built for. It takes up to a million base pairs, predicts thousands of regulatory outputs at once and gives you a mechanistic readout of what a variant is doing, how it’s disrupting regulation and at what stage.
Once you know that, you can intervene. If AlphaGenome shows that a condition comes down to a single nucleotide mutation destroying a splice site, CRISPR-Cas9 could fix it directly. CRISPR functions as molecular scissors that you can aim at an exact position in a genome to cut and replace the DNA. You’re fixing the typo in the original instruction manual. Every time the cell reads that gene going forward, it gets the right instructions. That’s as permanent of a fix as it gets. However, CRISPR isn’t always the best option. If the mutation is causing a splicing error, antisense oligonucleotides—short synthetic RNA fragments—can be sent into the cell where they latch onto the pre-mRNA at the broken splice site and physically block the spliceosome from making the wrong cut. The DNA stays untouched, the mutation stays where it is, but the bad instruction gets caught and corrected at the RNA level each time before it reaches a ribosome.
Nusinersen, a drug for spinal muscular atrophy, works on this principle and is being used on patients right now. If the mutation shuts a gene down entirely by destroying the enhancer driving its expression, gene therapy offers a third route. A working copy of the gene gets delivered into the cell using a viral vector. The mutation hasn’t gone anywhere but the cell now has a clean copy to read from. CRISPR fixes the DNA. Antisense oligonucleotides intercept the RNA. Gene therapy provides a backup. Which one you reach for depends on what the mutation is doing and where in the chain it’s causing damage, and that is exactly the information AlphaGenome gives you. Without it, you’re trying to fix a machine without knowing which part is broken.
In order to gain deeper understanding on the topic, I reached out to Professor Jeremy Dittman, M.D., Ph.D, of Wile-Cornell Medicine, who has far greater expertise in both biological science and the particular subject of what it’s like to work in a lab. Professor Dittman’s research specializes in molecular biology at the protein structure level and how mutations to protein structure cause functional differences in organisms. Specifically, his recent work has focused heavily on the SNARE complex, a set of proteins critical to neurotransmission in the neuron itself, specifically its axon terminal (the part of the neuron physically responsible for neurotransmission). Like almost all of the biomedical world, Professor Dittman has incorporated Alphafold predictions and AI generated protein structures into his work significantly since its release, using it to model the structure of normal proteins as well as using it to better understand how an amino acid substitution will affect the resulting structure of a given protein. During our interview, Prof. Dittman highlighted many interesting aspects of the Alphafold revolution and how it has affected science and the scientific community at large.
Firstly, Professor Dittman explained how Alphafold has changed work in the lab. Before Alphafold, a protein structure could take months to solve and was incredibly difficult. In order to figure out what mutations to research, and how those mutations could affect the protein’s structure, one had to consult whatever had already been written in medical literature and proceed with an experiment from there. This process itself could take months. However with Alphafold, generating a standard protein structure and its mutations can be done in as short as 15 minutes. The increase in speed and generative power is beyond exponential. This has allowed the lab to create new experiment ideas while also being able to verify their accuracy with a high level of probable truth via AF3 visualization. While AF3 results aren’t always correct, the mere ability to peer into these complex structures to know if we’re even chasing the right thing before the outset of wet lab work is powerful and removes a large migraine from molecular biologists minds.
After hearing about Alphafold’s success at CASP 2020, Dittman tried Alphafold for the first time after its release while writing a grant renewal. After pasting the sequence of the protein that the grant was centered on into Alphafold, new ideas about what to research were already flowing. Just by seeing this visualization, ideas and hypotheses were being written down within minutes. To Dittman, the greatest part about this new technology was the democratization of it all. With Google working as the primary funding and delivery system for the technology to the outside world, one could create all of these visualizations simply by looking up Alphafold 3 on Chrome. The tool is free and you can paste as many as 5,000 amino acids into it at one time for processing. “The fact that it is accessible to everyone has made us all into structural biologists.” said Dittman. “Alphafold is a golden opportunity for highschoolers to learn more about protein structure and mutations, but also gives anyone a chance to stumble onto answers and questions you didn’t initially ask.”
However, despite the visceral power and potential of Alphafold, Professor Dittman was careful to highlight the limitations and pitfalls of the current Alphafold infrastructure. One of the biggest misconceptions he feels has been floated about Alphafold is that it can/could replace research science entirely. In the professor’s opinion: “People say it’s going to replace experiments or that it could tell you the answer of an experiment, which it can’t. In order to verify that AF3 isn’t lying, you need to make an experiment to verify the protein structure of a molecule. It also cannot do any kind of modeling of an experiment that takes place over a course of time. The difference between Alphafold and experimentation is vast and not quantifiable by the variables AI can reliably work with at this point.”
In basic terms, Alphafold can model the structure of a protein attached to another protein, but it can’t tell you what those proteins will do when they interact. It cannot tell you where one part of the protein will go in relation to the other protein outside of the single stage of the interaction it chooses to model. This is a problem because it is equivalent to trying to understand gears in motion by taking a freeze frame of a single moment. Without the ability to model dynamic interactions and the fact that it cannot represent the ground truth of a molecule as a computer prediction, research scientists are far away from having their profession stolen by computer at this time.
Alphafold represents a true representation of the power of artificial intelligence when applied to novel concepts and a revolution for the world of protein biomedicine that will have powerful applications and effects on medical developments in the future.
Alphafold represents a true representation of the power of artificial intelligence when applied to novel concepts and a revolution for the world of protein biomedicine that will have powerful applications and effects on medical developments in the future.
