Protein sequence design
Introducing ProteinMPNN…
Common data representations for proteins in machine learning
Common data representations for proteins in machine learning
We covered protein structure prediction
ILVEPRTEINS…
ProteinMPNN predicts protein sequences given a backbone structure
Designing a sequence from a given backbone is often referred to as inverse folding
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
Encoder: Develops a sequence-independent representation of 3D structure
Decoder: Aggregates local neighborhood geometry. Places residues in positions that it sees as fitting best in context of geometry.
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
The MPNN architecture built for protein design by John Ingraham et al.
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
ProteinMPNN predicts protein sequences given a backbone structure
ProteinMPNN predicts protein sequences given a backbone structure
Trained on curated set of PDB structures:
ProteinMPNN predicts protein sequences given a backbone structure
Encoded input features are (mostly) comprised of just “edges” which are residue-residue connections on graph
Model builds and learns a sequence-independent representation of the structure based on these connections
Features - Edges
residue position neighbor position edge feature 1 edge feature 2 ... edge feature 128
1 7 ... ... ...
1 14 ... ... ...
1 22 ... ... ...
2 5 ... ... ...
2 19 ... ... ...
2 41 ... ... ...
...
68 12 ... ... ...
68 27 ... ... ...
68 63 ... ... ...
ProteinMPNN predicts protein sequences given a backbone structure
Decoder takes in encoded edge features + an encoded sequence to generate probabilities for amino acid identities at each position
It is the “translator” of the model. Takes the “lessons” from the structure representation, calculates probabilities of amino acids based on this geometry
ProteinMPNN predicts protein sequences given a backbone structure
ProteinMPNN changed linear autoregressive decoder to randomized autoregressive decoder scheme
ProteinMPNN predicts protein sequences given a backbone structure
ProteinMPNN changed left-to right autoregressive decoder to randomized decoding scheme
Quick Review
Is ProteinMPNN trained on physicochemical properties of amino acids such as charge, size, hydrophobicity?
Does ProteinMPNN use an energy function?
Are ProteinMPNN designs optimized for stability?
No
No
No (not explicitly)
ProteinMPNN is a (Relatively) Simple Model
Ingraham, et al., NeurIPS 2019
https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf
How Do You Assess Performance?
Experimental validation
Performance Assessment: Sequence Recovery
Native Seq: QLINTNGSWHIN
Design Seq: QLVNTQGSIHIN
9/12 residues “recovered” = 0.75 sequence recovery
Performance Assessment: Other Methods
A C D E F G H I K L M N P Q R S T V W Y
Native: GLU
Probability
Performance Assessment: Other Methods
ProteinMPNN performance
More nearby residues = more context to inform model = higher sequence recovery
AlphaFold performance on ProteinMPNN sequences
Increasing ProteinMPNN sampling temperature increases sequence diversity (and entropy)
Increasing ProteinMPNN sampling temperature increases sequence diversity (and entropy)
T
1.0
0.1
ProteinMPNN can design multi-chain pdbs
Figure from Simon Duerr
Fixed residues and chains are used as additional “context” when selecting which amino acid to put at each position
Limitations
Example Use Case 1: Enzyme Design
Used PMPNN to redesign myoglobin and TEV protease
Goal: Improve expression, stability, and activity
Only designed residues at positions away from active site with low conservation
Example Use Case 1: Enzyme Design
Most myoglobin redesigns were more soluble than wild type (13/20).
8/8 Evaluated designs had improved stability (5/8 better heme bind retention)
64/144 of TEV protease designs demonstrated higher activity than WT TEV
Some Important Design Considerations
Example Use Case 2: Making TMD Proteins Soluble
Example Use Case 2: Making TMD Proteins Soluble
Composition of output designs change dramatically from default PMPNN (way more polar residues)
Redesign of transmembrane proteins was fairly successful (70%) were soluble with smaller portion retaining function (not shown)
Example Use Case 3: Foundation of Mutation Prediction
Introduced a tool called ThermoMPNN where ProteinMPNN serves as a foundation for transfer learning approach.
Goal is to enable the prediction of mutations on the stability of a protein
Example Use Case 3: Foundation of Mutation Prediction
LigandMPNN
Inputs and Outputs
LigandMPNN
MGDIQVQVNIDDNGKNFDYTYTVTTESELQKVLNELMDYIKKQGAKRVRISITARTKKEAEKFAAILIKVFAELGYNDINVTFDGDTVTVEGQLEGGSLEHHHHHH
Predicted Sequence
Desired backbone topology + small molecule
Dauparas et al. (2023)
Additional atomic context improves local sequence recovery
Additional atomic context improves local sequence recovery
Additional atomic context improves local sequence recovery
LigandMPNN Overview
Other Sequence Design/Inverse Folding Methods
Slide credits
Nate Felbinger
Deniz Akpinaroglu