1 of 44

Protein sequence design

Introducing ProteinMPNN…

2 of 44

Common data representations for proteins in machine learning

Gao, Mahajan, Sulam & Gray Patterns 2020

https://doi.org/10.1016/j.patter.2020.100142

​

3 of 44

Common data representations for proteins in machine learning

Gao, Mahajan, Sulam & Gray Patterns 2020

https://doi.org/10.1016/j.patter.2020.100142

​

4 of 44

We covered protein structure prediction

ILVEPRTEINS…

5 of 44

ProteinMPNN predicts protein sequences given a backbone structure

Designing a sequence from a given backbone is often referred to as inverse folding

6 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

7 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

8 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

Encoder: Develops a sequence-independent representation of 3D structure

Decoder: Aggregates local neighborhood geometry. Places residues in positions that it sees as fitting best in context of geometry.

9 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

10 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

  • Takes X, Y , Z coordinates of backbone atoms
  • Uses radial basis function (RBF) to calculate distance to find neighboring residues
  • Connections between neighbors become the (edge) features
  • Implicitly invariant

11 of 44

The MPNN architecture built for protein design by John Ingraham et al.

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

12 of 44

ProteinMPNN predicts protein sequences given a backbone structure

13 of 44

ProteinMPNN predicts protein sequences given a backbone structure

Trained on curated set of PDB structures:

  • Released before 08-02-2021
  • X-ray or cryo-em better than 3.5A (no NMR structures)
  • Clustered at 30% sequence identity, non redundant

14 of 44

ProteinMPNN predicts protein sequences given a backbone structure

Encoded input features are (mostly) comprised of just “edges” which are residue-residue connections on graph

Model builds and learns a sequence-independent representation of the structure based on these connections

15 of 44

Features - Edges

residue position neighbor position edge feature 1 edge feature 2 ... edge feature 128

1 7 ... ... ...

1 14 ... ... ...

1 22 ... ... ...

2 5 ... ... ...

2 19 ... ... ...

2 41 ... ... ...

...

68 12 ... ... ...

68 27 ... ... ...

68 63 ... ... ...

16 of 44

ProteinMPNN predicts protein sequences given a backbone structure

Decoder takes in encoded edge features + an encoded sequence to generate probabilities for amino acid identities at each position

It is the “translator” of the model. Takes the “lessons” from the structure representation, calculates probabilities of amino acids based on this geometry

17 of 44

ProteinMPNN predicts protein sequences given a backbone structure

ProteinMPNN changed linear autoregressive decoder to randomized autoregressive decoder scheme

18 of 44

ProteinMPNN predicts protein sequences given a backbone structure

ProteinMPNN changed left-to right autoregressive decoder to randomized decoding scheme

19 of 44

Quick Review

Is ProteinMPNN trained on physicochemical properties of amino acids such as charge, size, hydrophobicity?

Does ProteinMPNN use an energy function?

Are ProteinMPNN designs optimized for stability?

No

No

No (not explicitly)

20 of 44

ProteinMPNN is a (Relatively) Simple Model

Ingraham, et al., NeurIPS 2019

https://proceedings.neurips.cc/paper_files/paper/2019/file/f3a4ff4839c56a5f460c88cce3666a2b-Paper.pdf

  • ProteinMPNN is only trained on backbones/sequence pairs of curated PDB structures
  • The sequences it designs are those it finds to be most probable to fit a given backbone
  • ProteinMPNN designs often improve solubility, stability, etc. but these are not goals of the model

21 of 44

How Do You Assess Performance?

Experimental validation

22 of 44

Performance Assessment: Sequence Recovery

  • One commonly used metric to assess the performance of sequence design models is sequence recovery
  • Sequence recovery measures the number of native amino acid identities that are also incorporated in the designed sequence
  • Higher sequence recovery != better sequence
  • However when used for benchmarking it can indicate how well sequence design models “understand” natural sequences

Native Seq: QLINTNGSWHIN

Design Seq: QLVNTQGSIHIN

9/12 residues “recovered” = 0.75 sequence recovery

23 of 44

Performance Assessment: Other Methods

  • Tokens of the decoder = Residue identities (which it pairs with embeddings from encoder)
  • Final output = Residue identities, but these must be derived from the logits

A C D E F G H I K L M N P Q R S T V W Y

Native: GLU

Probability

24 of 44

Performance Assessment: Other Methods

  • ProteinMPNN includes a pseudo-confidence metric by averaging the per-residue negative log probability (derived from logits) of the amino acids incorporated into the design output
  • Authors do not recommend relying on it to assess sequence
  • Instead you should model the design
    • Does the output model fold similarly to your input (RMSD/LDDT)?
    • Is the model output high confidence (confidence, pLDDT)?
  • Manual inspection of designs is also important. Are their ridiculous-looking repeats in the output? Prolines in secondary structure?

25 of 44

ProteinMPNN performance

More nearby residues = more context to inform model = higher sequence recovery

26 of 44

AlphaFold performance on ProteinMPNN sequences

27 of 44

Increasing ProteinMPNN sampling temperature increases sequence diversity (and entropy)

28 of 44

Increasing ProteinMPNN sampling temperature increases sequence diversity (and entropy)

T

1.0

0.1

29 of 44

ProteinMPNN can design multi-chain pdbs

Figure from Simon Duerr

Fixed residues and chains are used as additional “context” when selecting which amino acid to put at each position

30 of 44

Limitations

  • Trained on whole PDB so recognizes membrane structures
    • Soluble ProteinMPNN
  • Sensitive to quality of the backbone
    • Relax structures or produce templated AlphaFold models to get structures with “ideal” geometry
  • Does not account for non-protein molecules
    • LigandMPNN

31 of 44

Example Use Case 1: Enzyme Design

Used PMPNN to redesign myoglobin and TEV protease

Goal: Improve expression, stability, and activity

Only designed residues at positions away from active site with low conservation

32 of 44

Example Use Case 1: Enzyme Design

Most myoglobin redesigns were more soluble than wild type (13/20).

​

8/8 Evaluated designs had improved stability (5/8 better heme bind retention)

​

64/144 of TEV protease designs demonstrated higher activity than WT TEV

​

Some Important Design Considerations

​

  • Histidine and cysteine residues were excluded from designs
  • Moderate restriction of conserved residues worked best
  • Various sampling temperatures were used (0.1, 0.2, 0.3) did not appear to alter design success
  • Cyclical design process improved success rate (MPNN design -> inpaint -> MPNN design again)
  • MD simulations demonstrated difference in stabilizing design

​

33 of 44

Example Use Case 2: Making TMD Proteins Soluble

  • Goal here was to take complex insoluble proteins and produce soluble designs that maintained function
  • A lot more complexity to their design pipeline than just plugging sequences into ProteinMPNN
  • Used a fine-tuned model of ProteinMPNN. Essentially repeated same training protocol, but excluded transmembrane proteins from training data (SolubleMPNN).

​

34 of 44

Example Use Case 2: Making TMD Proteins Soluble

Composition of output designs change dramatically from default PMPNN (way more polar residues)

​

​

Redesign of transmembrane proteins was fairly successful (70%) were soluble with smaller portion retaining function (not shown)

​

​

35 of 44

Example Use Case 3: Foundation of Mutation Prediction

Introduced a tool called ThermoMPNN where ProteinMPNN serves as a foundation for transfer learning approach.

Goal is to enable the prediction of mutations on the stability of a protein

36 of 44

Example Use Case 3: Foundation of Mutation Prediction

37 of 44

LigandMPNN

38 of 44

Inputs and Outputs

LigandMPNN

  • sequences (protein + NA)
  • small molecules
  • covalent modifications

MGDIQVQVNIDDNGKNFDYTYTVTTESELQKVLNELMDYIKKQGAKRVRISITARTKKEAEKFAAILIKVFAELGYNDINVTFDGDTVTVEGQLEGGSLEHHHHHH

Predicted Sequence

Desired backbone topology + small molecule

Dauparas et al. (2023)

39 of 44

Additional atomic context improves local sequence recovery

40 of 44

Additional atomic context improves local sequence recovery

41 of 44

Additional atomic context improves local sequence recovery

42 of 44

LigandMPNN Overview

43 of 44

Other Sequence Design/Inverse Folding Methods

  • Many other sequence design tools exists, with alternative model architectures and capabilities:
    • Protein language model-based: ESM-IF1, ESM3
    • All-atom: FAMPNN
    • Others: Frame2Seq, piFold, BC-Design

44 of 44

Slide credits

Nate Felbinger

Deniz Akpinaroglu