Protein-folding AI has been one of the most celebrated scientific advances of recent years. Yet a quiet limitation has been building beneath the surface. The models that predict how proteins interact with potential drugs are only as good as the data they were trained on, and that data, it turns out, is far thinner than most people realize. A consortium of pharmaceutical companies has now demonstrated that unlocking proprietary molecular data, kept private for competitive reasons, can substantially close that gap.
The Data Vault Problem at the Heart of Drug Discovery
The Protein Data Bank, known as the PDB, is the open repository that made AlphaFold 2 possible. It contains more than 200,000 experimentally determined protein structures, and its scale was central to AlphaFold 2’s breakthrough accuracy, a contribution recognized with the 2024 Nobel Prize in Chemistry.
But the next generation of protein AI tools faces a different challenge. Models like AlphaFold 3 are designed not just to predict protein shapes, but to predict how proteins interact with other molecules, including candidate drugs. This is where the PDB falls short. According to Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals, the PDB contains only around 10,000 experimentally determined structures showing proteins interacting with drug-like molecules. That is a small number relative to the diversity of molecular interactions that drug discovery requires.
The rest of that data sits inside pharmaceutical companies. Generated through techniques such as X-ray crystallography and cryo-electron microscopy, these structures were never deposited in public databases because they relate to proprietary drug development programs. The total volume of this private data is unknown, but some estimates suggest it could exceed what the PDB holds. As John Karanicolas, head of computational drug discovery at AbbVie, put it: the data missing from the PDB is precisely the data present in pharmaceutical companies’ internal systems.
What Happens When You Pool the Private Data
To test whether that private data could improve AI performance, AbbVie, Astex, and several other drug companies formed a collaboration called the AI Structural Biology Network, or AISB. The group took OpenFold3, an open-source replication of AlphaFold 3 that had been trained only on public PDB data, and fine-tuned it on an additional 20,167 proprietary protein structures. These structures came from five companies and were shared in a way that kept each firm’s data private from the others.
The results were notable. When the AISB model was tested against 1,056 protein-ligand structures held back from training, it predicted more than half of them to a high level of accuracy. The publicly available version of OpenFold3, trained only on public data, achieved the same performance on just one-third of those structures. A competing open-source model called Boltz-2 reached around 40%.
Mohammed AlQuraishi, a computational biologist at Columbia University who participated in the effort, described the improvement as “a pretty big bump in performance.” Karanicolas noted that the AISB model also outperformed models trained on each individual company’s data alone, which underscores a key point: pooling information across organizations produces results that no single organization could achieve independently.
The study has been described in a blog post and has not yet been peer-reviewed. The model itself is not publicly available. The team plans to submit the work to a peer-reviewed journal.
Why This Matters Beyond the Lab
Here is what most coverage of protein AI misses: the bottleneck in drug discovery is not computational power, and it is not algorithmic sophistication. It is data. Specifically, it is the right kind of data, structured examples of proteins binding to drug-like molecules, that reflects the actual chemical diversity researchers encounter when developing new medicines.
The AISB experiment is a proof of concept for a broader idea: that scientific progress in AI-assisted drug discovery may depend less on building smarter models and more on building better data-sharing structures. The competitive instinct to keep molecular data private is understandable, but the AISB results suggest that collective data sharing, even among rivals, produces tools that benefit everyone.
This dynamic is not unique to pharmaceuticals. It echoes a pattern visible across AI development: the organizations that can aggregate diverse, high-quality training data gain capabilities that those working in isolation cannot match. In drug discovery, the stakes are particularly high. Improved protein-ligand prediction could accelerate the identification of viable drug candidates, reduce the cost of early-stage research, and ultimately shorten the path from laboratory hypothesis to clinical treatment.
Publicly funded efforts are also moving in this direction. A project called OpenBind, supported by up to £8 million in UK government funding, has already released hundreds of new protein structures, with thousands more planned.
In Short
AI protein-folding models have a data problem: public databases contain too few examples of proteins interacting with drug-like molecules. A consortium of pharmaceutical companies demonstrated that training on more than 20,000 proprietary structures, pooled across five firms, lifted prediction accuracy from roughly one-third to more than half of test cases. The finding reframes the challenge in AI-assisted drug discovery: the limiting factor is not the algorithm, it is access to the right data, and sharing that data, even among competitors, produces results no single organization can achieve alone.
Based on reporting from Nature: Machine Learning.