MIT researchers have developed PottsMPNN, a machine-learning framework intended to improve computational protein design while reducing the tendency to reproduce sequences found in nature.

Protein function depends on structure, while structure is determined by an amino acid sequence. Many design workflows first specify a structure and then use machine learning to generate sequences that might fold into it. Because multiple sequences can produce the same fold—and one flexible sequence can sometimes adopt different structures—models need to recognize that several answers may be viable.
“Native sequence recovery” has traditionally been used to evaluate these systems, but Amy E. Keating, head of MIT’s Department of Biology and senior author of a paper published in PNAS, says reproducing evolution’s selected sequence is not the best measure of design success. For a novel structure, there may be no native sequence for comparison. More relevant measures include whether generated sequences fold into the intended structure, how well the model represents the sequence-energy landscape, and whether it predicts mutation effects on stability.
Modeling sequence and stability
PottsMPNN incorporates physical principles governing protein structure and stability. It uses a pairwise distribution to represent interactions between all 20 amino acid options at pairs of positions, helping it model the sequence-energy landscape more accurately. The framework also uses evolutionarily related sequences during training to show how different sequences can form the same fold.
Foster Birnbaum, a graduate student and the paper’s lead author, initially explored adding “noise”—controlled variations to protein structures during training. This reduced the model’s tendency to mimic native sequences and broadened the range of structures for which it could generate sequences.

Although evolutionary information still introduces some reliance on native sequences, the researchers found that reducing dependence on those sequences improved structural compatibility and energy prediction, including for proteins that are novel.
Potential applications
Birnbaum said the framework could eventually be tuned for specific tasks, such as predicting the consequences of particular mutations. Keating said the work advances efforts to design useful new-to-nature proteins for diverse applications and provides a stronger basis for future progress.
Birnbaum also noted that the ability to design any protein could enable a “potentially scary amount of biological engineering,” while expressing optimism about advances in biology during this century.
The research paper is titled “Beyond native sequence recovery: Improved modeling of the sequence-energy landscape of protein structures.”



