1 of 20

MACHINE LEARNING POWERED HIGH CONTENT IMAGE ANALYSIS OF H&E BIOPSIES TO PREDICT OUTCOMES IN HPV+ OROPHARYNGEAL CANCERS��JONAS HUE, SELVAM THAVARAJ, LORENZO VESCHINI��KING’S COLLEGE LONDON, FACULTY OF DENTISTRY, ORAL & CRANIOFACIAL SCIENCES

2 of 20

Roadmap

  1. Introduction

​

  • Workflow

​

  • Potential applications

​

​

​

3 of 20

Background

  • Human papillomavirus (HPV) infection can induce carcinogenesis in the oropharynx with E6 & E7 viral oncoproteins

​

  • HPV+ oropharyngeal squamous cell carcinomas (OpSCCs) tend to have a favourable prognosis compared to HPV- OpSCCs
      • Treatment de-escalation to reduce side effects of chemo/radiotherapy

​

  • HPV+ OpSCCs are now considered as a separate entity from HPV- OpSCCs

​

  • However, a significant number (~15-20%) of patients still have poor outcomes
      • Potential biomarkers have been identified in this group of patient
      • Treatment de-escalation may not be indicated in these patients

​

​

​

4 of 20

Aims

  1. Image analysis workflow to quantify features of patient biopsies

​

  • Identification of prognostic features

​

  • Machine learning model to predict patient outcomes

​

​

​

5 of 20

Cohort

  • Retrospective test cohort
    • 29 patients with unfavourable outcomes (i.e. died of disease or recurrence within 5 years)
    • 29 with favourable outcomes (i.e. disease-free at 5 years)
  • Image Acquisition:
    • 10 representative fields at 100x magnification were selected from each biopsy aiming for ~2/3 tumour and 1/3 stroma

​

6 of 20

Workflow

1. H&E Images

3. FIJI & Stardist

2. QuPath

6. Data Analysis

7. Neural Network Model

4. CellProfiler

5. CellProfiler Analyst

Identification of 28 potentially prognostic features

  • Immune cells
  • Nuclear morphology
  • Texture/granularity features

7 of 20

H&E Stain Variation—QuPath

  • First round of preprocessing is done to account for H&E stain variation using a neural network pixel classifier in QuPath

​

​

​

​

8 of 20

Crowded Cells—Stardist

  • H&E sections are ~5microns thick, containing closely packed cells in more than 1 plane
  • A published, pre-trained, Stardist model for H&E was applied through an ImageJ plug-in to help with crowded cells
  • Stardist uses star-convex polygons and a convolutional neural network to predict where to draw boundaries between closely packed cells

​

​

​

9 of 20

Object-based Image Segmentation

  • Various cell types are identified as objects and specific measurements can be made of each individual object

​

​

​

​

​

10 of 20

Classifying Cells—CPA

  • Objects identified in CellProfiler can be classified with the help of supervised machine learning in CellProfiler Analyst (CPA)
  • Various models can be applied to classify various cell types or cell objects (eg. nucleoli)
  • We have attempted to classify tumour cells, lymphocytes, plasma cells, nucleoli

​

​

​

​

​

​

Tumour Cells

11 of 20

Classifying Cells—CPA

​

​

​

​

​

​

Plasma Cells

Nucleoli

Tumour Infiltrating Lymphocytes (TILs)

Tumour Cell Eccentricity

12 of 20

Data Analysis

  • Once the cells are classified, we can look at some of the following:
        • Quantity of immune cells
        • Spatial analysis and neighbours
        • Tumour cell morphology/ phenotypes

​

  • Measurements are exported to an SQLite database and further statistical analysis can be performed in R to identify factors with potential prognostic significance

​

​

​

​

​

13 of 20

Data Analysis

14 of 20

Prognostic Features

  • Features of Favourable Outcomes:
    • More TILs
    • More Plasma cells
    • Rounder and less eccentric tumour nuclei
    • Tumour cells packed closer
    • More nucleoli
    • Higher nuclear textural features
    • Higher nuclear granularity values
    • More homogenous tumour morphology
    • More homogenous tumour cell sizes
    • More homogenous spatial packing

15 of 20

Predictive Model—Neural Network

  • The 28 prognostic variables are fed into a neural network with 4 hidden layers (16, 8, 8, 4) to predict patient outcomes from their H&E biopsies
  • 567 images were split into 80% training and 20% test sets
    • Overall Accuracy of 78.76%,
    • Sensitivity of 72.13% for favourable outcomes
    • Specificity of 86.54% for favourable outcomes

​

16 of 20

Predictive Model—Neural Network

  • A k-fold cross-validation was also performed with 10 folds to validate the model

​

  • Average overall accuracy: 77.72%

​

  • Average Sensitivity: 77.17%

​

  • Average Specificity: 78.12%

​

​

​

17 of 20

Advantages of the Workflow

  • Uses routine H&E stains:
    • No additional biomarkers
  • Low cost:
    • All aspects of the workflow use free, open-source software
    • No additional staining required
    • No whole-slide scanning required
  • Low computational cost:
    • Each patient was analysed in under 2hrs with a standard laptop with 8GB RAM
  • Time efficient:
    • Does not delay diagnosis and initiation of cancer treatment
    • Does not require pathologist to manually analyse slides for various prognostic factors
  • Quantitative:
    • Not subject to operator bias or semi-quantitative manual assessment

​

​

​

​

18 of 20

Applying the Workflow to TILs in Breast Cancer

  • A screenshot from one of the biopsies from The Cancer Genome Atlas, run through the workflow (https://mathbiol.github.io/tcgatil/)
  • The pipeline is fairly modular and it is possible to only segment and identify TILs, allowing it to be applied to other cancer types

​

​

​

​

19 of 20

Other Potential Applications

  • Single-cell segmentation and classification allows us to obtain spatial maps of the various cell types

​

​

​

20 of 20

THANK YOU!