Predicting Potential Drug Targets Using Tensor Factorisation and Knowledge Graph Embeddings
Cheng Ye, Rowan Swiers, Stephen Bonner and Ian Barrett
Data Sciences and Quantitative Biology, Discovery Sciences
R&D, AstraZeneca
Cambridge, UK
23/09/2021
Please ensure you have added the correct confidentiality statement.
Type the correct statement: �‘Company Restricted’ or ‘Strictly Confidential’ . No statement is needed �for public information.
!
Drug Target Identification - what is the problem and why is it important
Drug discovery and development process is a long and expensive one, costing over $1 billion on average per drug and taking 10-15 years
Selecting an efficacious drug target is the first and arguably the most important step, nevertheless, more than half of clinical trials fail due to lack of efficacy in pharmaceutical industry
2
Cook et al., Nature reviews Drug Discovery, 2014
https://m.dragonrest.net/optimizing-drug-development-with-model-informed-decision-making
A new challenge: Data-Driven Target Identification
3
Traditional approaches to target identification
Discovering novel insights into a disease mechanism from research publications and designing a drug molecule to activate or inhibit the target to cease or reverse the disease effects
Repositioning existing drugs to find unanticipated but beneficial effects against other diseases
These processes have generated a large number of therapeutic hypotheses which are further supported by an enormous amount of both structured and unstructured evidence data
Predicting disease targets using matrix factorisation
4
Natarajan et al., Bioinformatics, 2014
Side Information
Matrix Factorisation
Upgrading the framework with �tensor factorisation and knowledge graph embeddings
5
Why using tensor instead of matrix
A matrix cannot represent the full heterogeneous evidence data supporting gene-disease associations and the clinical outcomes
6
Why using knowledge graph embeddings as side information
A knowledge graph is a model of a knowledge domain that interconnects information and captures the meaning of those interconnections
Knowledge graphs allow feature representations (embeddings) for different entities to be automatically learned
Advantages over manually engineered side information
7
Our new framework for predicting the clinical outcome of a target-disease pair
8
Model details:
Experimental setup and results
9
Datasets overview
Three-dimensional tensor (density = 3.19%)
Drug discovery-oriented knowledge graph (Hetionet, https://het.io/)
10
Evaluation strategies
11
Tensor factorisation outperforms baseline ML models
12
Incorporating knowledge graph embeddings as side information increases the performance
13
Conclusions and future work
Developed a novel tensor factorisation framework to identify clinically promising targets for diseases – will help reduce the time and expense of the drug pre-discovery stage
Our main contributions
Future research directions
14
Acknowledgements
15
16
Thank you.
17
Confidentiality Notice
This file is private and may contain confidential and proprietary information. If you have received this file in error, please notify us and remove it from your system and note that you must not copy, distribute or take any action in reliance on it. Any unauthorized use or disclosure of the contents of this file is not permitted and may be unlawful. AstraZeneca PLC, 1 Francis Crick Avenue, Cambridge Biomedical Campus, Cambridge, CB2 0AA, UK, T: +44(0)203 749 5000, www.astrazeneca.com