1 of 17

Predicting Potential Drug Targets Using Tensor Factorisation and Knowledge Graph Embeddings

Cheng Ye, Rowan Swiers, Stephen Bonner and Ian Barrett

Data Sciences and Quantitative Biology, Discovery Sciences

R&D, AstraZeneca

Cambridge, UK

23/09/2021

Please ensure you have added the correct confidentiality statement.

Type the correct statement: �‘Company Restricted’ orStrictly Confidential’ . No statement is needed �for public information.

!

2 of 17

Drug Target Identification - what is the problem and why is it important

Drug discovery and development process is a long and expensive one, costing over $1 billion on average per drug and taking 10-15 years

Selecting an efficacious drug target is the first and arguably the most important step, nevertheless, more than half of clinical trials fail due to lack of efficacy in pharmaceutical industry

2

Cook et al., Nature reviews Drug Discovery, 2014

https://m.dragonrest.net/optimizing-drug-development-with-model-informed-decision-making

3 of 17

A new challenge: Data-Driven Target Identification

3

Traditional approaches to target identification

Discovering novel insights into a disease mechanism from research publications and designing a drug molecule to activate or inhibit the target to cease or reverse the disease effects

Repositioning existing drugs to find unanticipated but beneficial effects against other diseases

These processes have generated a large number of therapeutic hypotheses which are further supported by an enormous amount of both structured and unstructured evidence data

4 of 17

Predicting disease targets using matrix factorisation

4

Natarajan et al., Bioinformatics, 2014

Side Information

Matrix Factorisation

 

5 of 17

Upgrading the framework with �tensor factorisation and knowledge graph embeddings

5

6 of 17

Why using tensor instead of matrix

A matrix cannot represent the full heterogeneous evidence data supporting gene-disease associations and the clinical outcomes

6

7 of 17

Why using knowledge graph embeddings as side information

A knowledge graph is a model of a knowledge domain that interconnects information and captures the meaning of those interconnections

Knowledge graphs allow feature representations (embeddings) for different entities to be automatically learned

Advantages over manually engineered side information

    • Data automatically incorporated – no need to hand-pick relevant data
    • Features effectively capture the roles of genes and diseases in the broader biomedical field

7

8 of 17

Our new framework for predicting the clinical outcome of a target-disease pair

8

Model details:

  • Embedding dimensionality = 32
  • Neural network size = (256, 64)
  • ReLU activation
  • Adagrad optimiser
  • MSE loss function

9 of 17

Experimental setup and results

9

10 of 17

Datasets overview

Three-dimensional tensor (density = 3.19%)

    • 1,048 gene targets
    • 860 diseases
    • 230,011 gene-disease evidence attributes and clinical outcomes

Drug discovery-oriented knowledge graph (Hetionet, https://het.io/)

    • 47,031 biological entities of 11 types (Gene, Disease, Compound, Pathway, Anatomy etc.)
    • 2,250,197 relationships of 24 types (Disease-associates-Gene, Compound-treats-Disease, Gene-participates-Pathway, etc.)

10

11 of 17

Evaluation strategies

  1. Random split of all gene-disease pairs (with clinical outcomes)
    • Standard cross-validation procedure
    • Introduces bias (pairs that share the same gene or the same disease could end up in different folds)

  • Randomly split genes first, diseases are then assigned into different folds accordingly

  • Split diseases into different categories, genes are then assigned into different folds accordingly

11

12 of 17

Tensor factorisation outperforms baseline ML models

12

13 of 17

Incorporating knowledge graph embeddings as side information increases the performance

13

14 of 17

Conclusions and future work

Developed a novel tensor factorisation framework to identify clinically promising targets for diseases – will help reduce the time and expense of the drug pre-discovery stage

Our main contributions

    • Built a tensor to store gene-disease evidence and clinical outcomes
    • Established a neural network-based training paradigm of tensor factorisation
    • Used drug discovery-oriented knowledge graph embeddings as the side information

Future research directions

    • Exploit data negativeness – multiple reasons could be responsible for the failure of a clinical trial: drug safety (toxicology or clinical safety), drug efficacy (failure to achieve sufficient efficacy), pharmacokinetics/pharmacodynamics (PK/PD), commercial benefits, company strategy, etc.
    • Manually set (or learn) the weights of positive samples – some successful clinical trials are more trustworthy

14

15 of 17

Acknowledgements

15

16 of 17

16

Thank you.

17 of 17

17

Confidentiality Notice

This file is private and may contain confidential and proprietary information. If you have received this file in error, please notify us and remove it from your system and note that you must not copy, distribute or take any action in reliance on it. Any unauthorized use or disclosure of the contents of this file is not permitted and may be unlawful. AstraZeneca PLC, 1 Francis Crick Avenue, Cambridge Biomedical Campus, Cambridge, CB2 0AA, UK, T: +44(0)203 749 5000, www.astrazeneca.com