PottsMPNN aims to expand AI-designed protein possibilities

PottsMPNN aims to expand AI-designed protein possibilities

MIT researchers have developed PottsMPNN, a machine-learning framework intended to improve computational protein design while reducing the tendency to reproduce sequences found in nature.

Diagrams of five protein structures, one of which (green) is natural while the others were generated from designed sequences. Each structure has a halo of its amino acid sequence.

Protein function depends on structure, while structure is determined by an amino acid sequence. Many design workflows first specify a structure and then use machine learning to generate sequences that might fold into it. Because multiple sequences can produce the same fold—and one flexible sequence can sometimes adopt different structures—models need to recognize that several answers may be viable.

“Native sequence recovery” has traditionally been used to evaluate these systems, but Amy E. Keating, head of MIT’s Department of Biology and senior author of a paper published in PNAS, says reproducing evolution’s selected sequence is not the best measure of design success. For a novel structure, there may be no native sequence for comparison. More relevant measures include whether generated sequences fold into the intended structure, how well the model represents the sequence-energy landscape, and whether it predicts mutation effects on stability.

Modeling sequence and stability

PottsMPNN incorporates physical principles governing protein structure and stability. It uses a pairwise distribution to represent interactions between all 20 amino acid options at pairs of positions, helping it model the sequence-energy landscape more accurately. The framework also uses evolutionarily related sequences during training to show how different sequences can form the same fold.

Foster Birnbaum, a graduate student and the paper’s lead author, initially explored adding “noise”—controlled variations to protein structures during training. This reduced the model’s tendency to mimic native sequences and broadened the range of structures for which it could generate sequences.

Article image

Although evolutionary information still introduces some reliance on native sequences, the researchers found that reducing dependence on those sequences improved structural compatibility and energy prediction, including for proteins that are novel.

Potential applications

Birnbaum said the framework could eventually be tuned for specific tasks, such as predicting the consequences of particular mutations. Keating said the work advances efforts to design useful new-to-nature proteins for diverse applications and provides a stronger basis for future progress.

Birnbaum also noted that the ability to design any protein could enable a “potentially scary amount of biological engineering,” while expressing optimism about advances in biology during this century.

The research paper is titled “Beyond native sequence recovery: Improved modeling of the sequence-energy landscape of protein structures.”

Connor Coley standing in the lab
Circling arrows depict the looping process and say: “Test; Learn; Generate; Review; Build…” with stylized proteins in center.
Three ribbon diagrams of proteins in a 3D format, each in different colors
Using a computer modeling approach that they developed, MIT biologists identified three different proteins that can bind selectively to each of three similar targets, all members of the Bcl-2 family of proteins.

Compartilhar este artigo