Scramble Life Sciences All articles
Gene Therapy & Genomic Medicine

From Sequence to Structure: How AI Is Rewriting the Rules of Drug Discovery

Scramble Life Sciences
From Sequence to Structure: How AI Is Rewriting the Rules of Drug Discovery

Photo: Mary Rose G. Tandang-Silvas, Takako Fukuda, Chisato Fukuda, Krisna Prak, Cerrone Cabanos, Aiko Kimura, Takafumi Itoh, Bunzo Mikami, Shigeru Utsumi, Nobuyuki Maruyama,Conservation and divergence on plant seed 11S globulins based on crystal structure

For most of modern biology's history, understanding a protein meant enduring years of painstaking laboratory work. Crystallography, cryo-electron microscopy, nuclear magnetic resonance spectroscopy—each method offered a narrow, expensive window into the three-dimensional architecture that determines what a protein does and, crucially, how it might be targeted by a drug. The process was slow by design, governed by chemistry and physics that did not bend to ambition or funding cycles.

That paradigm has now been fundamentally disrupted. When DeepMind's AlphaFold system achieved near-experimental accuracy in predicting protein structures during the 2020 Critical Assessment of Protein Structure Prediction competition, the structural biology community experienced something between vindication and vertigo. A computational model trained on evolutionary sequence data had, in many respects, solved a problem that had resisted the field's best efforts for more than five decades.

The question now is not whether AI-driven protein structure prediction works. It demonstrably does. The question is how quickly medicine can absorb the implications.

Reframing the Central Problem

Proteins are the molecular machinery of life—enzymes that catalyze reactions, receptors that receive signals, structural scaffolds that give cells their shape. Their function is inseparable from their three-dimensional form, which is itself determined by the sequence of amino acids encoded in a gene. This relationship—sequence to structure to function—is so foundational that it underpins virtually every discipline in biomedicine.

For drug developers, the structure of a target protein is the map by which a therapeutic is designed. Without it, researchers are navigating blind, relying on trial-and-error screening of compound libraries. With it, they can engineer molecules that fit a binding pocket the way a key fits a lock.

AlphaFold, and the open database of more than 200 million predicted protein structures that the European Bioinformatics Institute subsequently released, effectively handed researchers that map for the vast majority of known proteins—including thousands for which no experimental structure had ever been determined.

"We went from having structural information on perhaps a third of the human proteome to having reasonable predictions for essentially all of it," noted one computational biologist at a leading academic medical center in Boston. "That is not an incremental advance. That is a phase transition."

Rare Disease: The Immediate Beneficiary

Perhaps nowhere is the practical impact more immediate than in rare and orphan disease research. Many rare conditions are caused by mutations in a single gene that produces a misfolded or dysfunctional protein. Because patient populations are small and commercial incentives are limited, these diseases have historically attracted modest investment and even more modest structural data.

AlphaFold changes the calculus. Researchers at several US academic institutions and biotech firms have begun using predicted structures to identify candidate small molecules for conditions that previously lacked any tractable drug target. Startups focused on rare metabolic disorders, inherited neuropathies, and congenital enzyme deficiencies are now building early-stage pipelines on a structural foundation that would have taken a decade to assemble through conventional means.

The Broad Institute, the NIH's National Center for Advancing Translational Sciences, and a growing cohort of venture-backed biotechs in the Boston-Cambridge and San Francisco Bay Area corridors have all publicly cited AlphaFold-derived insights as material contributors to ongoing programs. For patients with conditions that have long been considered untreatable simply because the underlying biology was too poorly understood to intervene, this represents a meaningful acceleration.

Structure-Based Drug Design Grows Up

Structure-based drug design is not a new discipline. Pharmaceutical companies have practiced it for decades, and blockbuster HIV protease inhibitors developed in the 1990s stand as early proof that designing drugs around a known protein structure can yield transformative medicines.

What AI prediction adds is scale and speed. Where a structural biology team might previously have spent eighteen months determining the structure of a single target, computational tools now deliver usable predictions in hours. This does not eliminate the need for experimental validation—predicted structures carry uncertainties that must be tested, particularly in disordered regions and protein-protein interaction interfaces—but it dramatically compresses the front end of the discovery process.

Several large pharmaceutical companies, including those with significant US research operations, have integrated AlphaFold predictions into their standard target assessment workflows. At least two publicly traded biotechs have disclosed that their lead programs were initiated in part because AI structural predictions revealed previously unappreciated binding sites on targets that had been considered undruggable.

The term "undruggable" itself is now under revision. A protein once dismissed because its structure was unknown or appeared to lack obvious pockets may, upon computational examination, reveal opportunities that were simply invisible before.

Personalized Medicine and the Variant Problem

Beyond fixed targets, AI protein prediction is beginning to intersect with personalized medicine in ways that could reshape how clinicians interpret genetic data. The human genome harbors millions of variants of uncertain significance—sequence differences whose functional consequences are unknown. For a physician counseling a patient about a genetic test result, these ambiguous variants are a persistent source of uncertainty.

Computational tools that can predict how a specific amino acid substitution alters a protein's structure—and, by extension, its function—offer a path toward resolving some of that uncertainty. Researchers are now training models not only to predict wild-type structures but to assess the structural consequences of individual mutations, effectively building a library of variant-structure-function relationships.

This work is still early. Translating a predicted structural perturbation into a confident clinical interpretation requires layers of experimental and clinical validation that take years to accumulate. But the trajectory is clear, and the potential to move variant interpretation from probabilistic inference to mechanistic understanding is one of the more consequential long-term applications of the technology.

The Limits of the Revolution

Intellectual honesty requires acknowledging what AI protein prediction does not do. It does not model protein dynamics—the constant motion and conformational changes that govern how proteins behave in a cellular environment. It struggles with intrinsically disordered proteins, which comprise a substantial fraction of the proteome and include many disease-relevant targets. And it does not yet reliably predict how proteins interact with each other in complexes, though newer models including AlphaFold-Multimer are beginning to address this gap.

Perhaps most importantly, a predicted structure is a hypothesis, not a fact. The history of drug discovery is littered with programs that foundered because a computationally attractive target proved far more complex in biology than in silico. The new tools reduce the cost of generating hypotheses; they do not reduce the cost of being wrong.

A New Foundation for Biomedical Science

Despite its limitations, the structural prediction revolution represents a genuine inflection point. The biological sciences are increasingly defined by the interplay between experimental and computational approaches, and protein structure prediction has become one of the most vivid illustrations of what that interplay can accomplish.

For the researchers, clinicians, and patients navigating the long road from molecular discovery to approved therapy, AlphaFold and the ecosystem it has catalyzed do not promise miracles on a fixed timeline. What they offer is a more complete map of the biological territory—one that makes the journey, if not shorter, at least less likely to end in a dead end that could have been foreseen.

All Articles

Related Articles

The CAR-T Paradox: A Cure in the Clinic, a Crisis in the Supply Chain

The CAR-T Paradox: A Cure in the Clinic, a Crisis in the Supply Chain

The Unfinished Business of CRISPR: Gene Editing's Promise Meets Its Most Stubborn Adversary — Biology Itself

The Unfinished Business of CRISPR: Gene Editing's Promise Meets Its Most Stubborn Adversary — Biology Itself

Beyond the Booster: How Upstart Biotechs Are Engineering the Next Era of mRNA Medicine

Beyond the Booster: How Upstart Biotechs Are Engineering the Next Era of mRNA Medicine