1 of 28

Graph convolutional neural networks: Part 2

Oct 28th, 2021

BMI 826-23 Computational Network Biology�Fall 2021

Anthony Gitter

https://compnetbiocourse.discovery.wisc.edu

2 of 28

Topics in this section

  • Representation learning on graphs
  • Graph neural networks
  • Generative graph models

3 of 28

Goals for today

  • Variations on graph convolutions
    • Edge features
    • Alternative ways to handle variable node degree
    • Graph attention
    • Pooling
  • Biological applications
    • Protein function prediction
    • Epigenomics
    • Drug side effects
    • Drug discovery

4 of 28

Basic graph convolutional neural network

  • Assumptions made in graph convolution layer we saw in Part 1
    • No edge features
    • All neighbors should be treated in the same way (shared weights)
    • All neighbors are equally important
    • Do not need to aggregate information until the end

5 of 28

Edge features in graph convolutions

  • Recall protein graph had edge features
  • Minor modification makes it easy to update in first graph convolution layer
  • Unlike node features, harder to update in later layers

Image from Fout et al. NIPS 2017

 

 

6 of 28

Unique neighbor weights in graph convolutions

  • Have been using the same weight vector for all neighbors
  • Can generalize if we can sort neighbors
    • In protein interface prediction, can sort by residue-residue distance

 

 

7 of 28

Protein interface prediction results

  • More complex network architectures did not have a big impact

Graph architecture

Optimal graph layers

AUC

None

1

0.812

Node weights

3

0.891

Node and edge weights

2

0.898

Unique node weights and edge weights

3

0.891

8 of 28

GraphSAGE: alternative approach for graphs with variable degree

  •  

9 of 28

GraphSAGE: alternative approach for graphs with variable degree

Image from Hamilton et al. NIPS 2017

Sampling naturally handles variable degree

Defined multiple aggregation strategies in addition to mean

10 of 28

Which neighbors are important

  • Graph convolutional neural networks give each neighbor the same influence
  • What if some neighbors are more important?
  • How do we know which neighbors are important in advance?

  • Attention is a deep learning approach to address this
    • Originally developed for sequence models
    • Key idea: learn which neighbors are important

11 of 28

Graph Attention Networks

  •  

5

1

3

2

4

 

5

1

3

2

4

 

 

 

 

12 of 28

Calculating attention in graphs

Image from Veličković et al. ICLR 2018

 

Use a simple neural network to calculate attention

13 of 28

Calculating attention in graphs

Image from Veličković et al. ICLR 2018

 

Can have parallel kinds of attention (multiple heads)

Concatenate or aggregate them

 

14 of 28

Calculating attention in graphs

Need to normalize attention over neighbors

Use softmax function

 

 

 

 

Before normalization

After normalization

 

 

 

 

 

15 of 28

Graph Attention Networks: putting it all together

Simplified notation from Zhang et al. in DGL docs

Transform current representation

Calculate edge attention as function of node representations

Normalize edge attentions

Use edge attentions in node updates

Concatenation

16 of 28

Pooling

  • In 1D or 2D convolutional neural networks, pooling is common
  • Search for patterns in small regions, then aggregate, repeat

Initially recognize holes in the ball

Then recognize configurations of holes

17 of 28

Pooling

  • Pooling is important in graphs as well but challenging
    • No pre-defined grid structure
    • Can make computation more efficient
    • Could help generalize across graphs

  • Consider how to learn the pooling
    • Self-attention graph pooling
    • Lee et al. 2019 arXiv:1904.08082

18 of 28

Self-attention graph pooling

Image from Lee et al. 2019 arXiv:1904.08082

Compute attention with graph convolution

Select top k fraction of the nodes to keep

Use attention mask to shrink the graph

19 of 28

Biological applications

  • Initially popular in chemistry / biochemistry
  • Gaining traction in many other areas
  • Biological network analysis with deep learning

  • Key considerations:
    • Unsupervised / supervised
    • Node / edge / graph predictions
    • What varies? Node features, graph structure, etc.

20 of 28

Predicting quantitative protein function

  • Input: Wild type protein structure, amino acid sequence
  • Output: Functional activity score
  • Gelman et al. 2020 bioRxiv 353946

21 of 28

Predicting quantitative protein function

  • Importance of control experiments
  • Does graph neural network performance change with alternative graphs?

  • For these datasets, can often still get good performance with the control graphs
  • Requires fully connected layer after the graph layers

Gelman et al. 2020 bioRxiv 353946

22 of 28

Predicting protein function

  • Input: Protein structure, residue (node) features
  • Output: Gene Ontology labels
  • Gligorijevic et al. 2021 Nature Communications

23 of 28

Predicting epigenetic state

  • Input: DNA sequence, Hi-C contact maps
  • Output: TF binding, histone modifications, accessibility
  • Lanchantin and Qi 2020 Bioinformatics

24 of 28

Predicting polypharmacy side effects

  • Input: 3 types of protein and drug interactions, node features
  • Output: Types of side effects for drug-drug pair
  • Zitnik et al. 2018 Bioinformatics

25 of 28

Predicting drug properties

  • Input: Chemical structure
  • Output: Chemical’s ability to inhibit a drug target
  • Chen et al. 2018 Drug Discovery Today

Does this chemical bind an important protein domain and affect bioactivity?

26 of 28

Predicting drug properties

  • IDG-DREAM Drug-Kinase Binding Prediction Challenge
  • Predict binding affinity between 25-70 compounds and ~200 kinases
  • One top performer used graph convolutional neural networks
    • Park et al. 2019 DMIS_DK Submission
    • Tested multiple graph architectures in ensemble
    • Graph Attention Networks most important in ensemble but random forest also competetive
    • 78 node features from RDKit: atom type, degree, bound Hs, valence, aromatic

27 of 28

Conclusions

  • Graph convolutional networks can be applied to many of the other biological tasks in this course
  • Can be advantages to coupling node representation with final prediction
  • Very flexible formulation, can tweak almost anything
  • Many hyperparameters, non-trivial to train
  • Critical to benchmark against simpler models

28 of 28

Other active research areas

  • What graph structures can classic graph neural networks not distinguish?
    • Xu et al. 2018 arXiv:1810.00826
  • How to make graph neural networks equivariant to rotations, translations, reflections and permutations
    • Garcia Satorras et al. 2021 arXiv:2102.09844
  • Contrastive learning for unsupervised graph representation learning
    • Sun et al. 2020 ICLR 2020
  • Pre-training graph neural networks
    • Hu et al. 2020 ICLR 2020