For decades, understanding how viruses work at the molecular level has required painstaking laboratory work, expensive equipment, and years of effort. A significant portion of that challenge comes down to a deceptively simple question: what do viral proteins actually look like in three dimensions, and how do they interact with each other? A major expansion of the AlphaFold Protein Structure Database is now addressing that question at a scale that was previously unimaginable.
Researchers have added more than 8,000 virus protein dimers, which are pairs of interacting protein molecules, to the AlphaFold database. These structures come from 23 virus families, all of which include members capable of infecting humans. The addition is part of a newly launched pandemic preparedness portal within the database, which is maintained by the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI) in Hinxton, UK, and is freely available to its more than three million users worldwide.
Why Viruses Were a Blind Spot in the First Place
The AlphaFold database already holds predicted structures for the vast majority of known proteins on Earth. Viruses, however, have been a persistent gap. Joe Grove, a molecular virologist at the University of Glasgow, describes them plainly as a “blind spot.” The reason is partly biological and partly technical.
Many viruses, including flaviviruses such as Zika and dengue, replicate by first producing a single large molecule called a polyprotein. That polyprotein is then cut into individual functional proteins. The problem is that a virus’s genetic sequence does not always make it obvious where one protein ends and the next begins. This ambiguity leads to incomplete or inaccurate structure predictions, which is why high-quality entries for many viral proteins have been lacking until now.
To address this, researchers at the Swiss Institute of Bioinformatics in Geneva identified accurate sequences for thousands of viral proteins derived from polyproteins. The broader team then analyzed sequences from 41,774 proteins across approximately 2,800 viruses, including those responsible for mpox, measles, and hepatitis B.
What Was Added, and What Was Left Out
Using AlphaFold2, the research team generated predictions for 40,746 homodimers (pairs of identical protein strands) and nearly 1.7 million heterodimers (pairs of different molecules). Not all of these made it into the public database. Only 2,749 of the homodimers and 5,279 of the heterodimers were judged accurate enough for inclusion, though all predictions were made publicly available regardless.
The distinction matters. Being included in the curated database signals a level of confidence that makes the data more immediately useful for researchers designing experiments or developing treatments. The broader set of predictions, while accessible, requires more caution in interpretation.
There are also known limitations to what these predictions capture. The current structures do not include the sugar molecules that coat many viral proteins and help them evade the immune system. Larger protein complexes, known as trimers or higher-order assemblies, are also absent from the database in many cases. The spike protein that allows SARS-CoV-2 and other coronaviruses to infect host cells, for example, is made up of three identical proteins. So is HIV’s envelope-entry protein. Dimer-level predictions for these structures were not accurate enough to include. Sameer Velankar, a bioinformatician at EMBL-EBI, acknowledges that adding trimer predictions is a logical next step, though larger complexes present an additional challenge: it is not always clear how many copies of each component they contain.
What This Means Beyond the Lab
Here is what most coverage of this development misses: the significance is not just about having more data. It is about changing the speed and accessibility of a foundational step in virology.
Historically, determining a protein’s three-dimensional structure required either X-ray crystallography or cryo-electron microscopy, both of which are time-consuming, resource-intensive, and dependent on physical samples. AlphaFold2 generates predictions computationally, from sequence data alone. That shift means researchers in institutions without access to expensive imaging infrastructure can now engage with structural data that was previously out of reach.
The pandemic preparedness framing is also worth taking seriously. When a novel virus emerges, one of the first scientific priorities is understanding its proteins: how they function, how they interact, and where they might be vulnerable to drugs or antibodies. Having a curated, freely accessible database of viral protein structures does not replace laboratory validation, but it compresses the time between “a new pathogen has appeared” and “here are the molecular targets worth investigating.”
Grove himself frames the next phase in practical terms: the value of the data will become clear as research teams begin working with it, testing predictions against experimental results, and identifying which structures hold up under scrutiny.
In Short
The AlphaFold database has added more than 8,000 viral protein pair structures from 23 human-infecting virus families, alongside a new pandemic preparedness portal. The expansion addresses a longstanding gap caused by the complexity of viral polyproteins. Limitations remain, including the absence of sugar molecule coatings and the exclusion of larger protein complexes. What the addition represents, at its core, is a meaningful reduction in the barrier between a viral outbreak and a molecular-level understanding of what researchers are dealing with.
Based on reporting from Nature: Machine Learning.