1 of 20

MACHINE LEARNING POWERED HIGH CONTENT IMAGE ANALYSIS OF H&E BIOPSIES TO PREDICT OUTCOMES IN HPV+ OROPHARYNGEAL CANCERS��JONAS HUE, SELVAM THAVARAJ, LORENZO VESCHINI��KING’S COLLEGE LONDON, FACULTY OF DENTISTRY, ORAL & CRANIOFACIAL SCIENCES

2 of 20

Roadmap

  1. Introduction

  • Workflow

  • Potential applications

3 of 20

Background

  • Human papillomavirus (HPV) infection can induce carcinogenesis in the oropharynx with E6 & E7 viral oncoproteins

  • HPV+ oropharyngeal squamous cell carcinomas (OpSCCs) tend to have a favourable prognosis compared to HPV- OpSCCs
      • Treatment de-escalation to reduce side effects of chemo/radiotherapy

  • HPV+ OpSCCs are now considered as a separate entity from HPV- OpSCCs

  • However, a significant number (~15-20%) of patients still have poor outcomes
      • Potential biomarkers have been identified in this group of patient
      • Treatment de-escalation may not be indicated in these patients

4 of 20

Aims

  1. Image analysis workflow to quantify features of patient biopsies

  • Identification of prognostic features

  • Machine learning model to predict patient outcomes

5 of 20

Cohort

  • Retrospective test cohort
    • 29 patients with unfavourable outcomes (i.e. died of disease or recurrence within 5 years)
    • 29 with favourable outcomes (i.e. disease-free at 5 years)
  • Image Acquisition:
    • 10 representative fields at 100x magnification were selected from each biopsy aiming for ~2/3 tumour and 1/3 stroma

6 of 20

Workflow

1. H&E Images

3. FIJI & Stardist

2. QuPath

6. Data Analysis

7. Neural Network Model

4. CellProfiler

5. CellProfiler Analyst

Identification of 28 potentially prognostic features

  • Immune cells
  • Nuclear morphology
  • Texture/granularity features

7 of 20

H&E Stain Variation—QuPath

  • First round of preprocessing is done to account for H&E stain variation using a neural network pixel classifier in QuPath

8 of 20

Crowded Cells—Stardist

  • H&E sections are ~5microns thick, containing closely packed cells in more than 1 plane
  • A published, pre-trained, Stardist model for H&E was applied through an ImageJ plug-in to help with crowded cells
  • Stardist uses star-convex polygons and a convolutional neural network to predict where to draw boundaries between closely packed cells

9 of 20

Object-based Image Segmentation

  • Various cell types are identified as objects and specific measurements can be made of each individual object

10 of 20

Classifying Cells—CPA

  • Objects identified in CellProfiler can be classified with the help of supervised machine learning in CellProfiler Analyst (CPA)
  • Various models can be applied to classify various cell types or cell objects (eg. nucleoli)
  • We have attempted to classify tumour cells, lymphocytes, plasma cells, nucleoli

Tumour Cells

11 of 20

Classifying Cells—CPA

Plasma Cells

Nucleoli

Tumour Infiltrating Lymphocytes (TILs)

Tumour Cell Eccentricity

12 of 20

Data Analysis

  • Once the cells are classified, we can look at some of the following:
        • Quantity of immune cells
        • Spatial analysis and neighbours
        • Tumour cell morphology/ phenotypes

  • Measurements are exported to an SQLite database and further statistical analysis can be performed in R to identify factors with potential prognostic significance

13 of 20

Data Analysis

14 of 20

Prognostic Features

  • Features of Favourable Outcomes:
    • More TILs
    • More Plasma cells
    • Rounder and less eccentric tumour nuclei
    • Tumour cells packed closer
    • More nucleoli
    • Higher nuclear textural features
    • Higher nuclear granularity values
    • More homogenous tumour morphology
    • More homogenous tumour cell sizes
    • More homogenous spatial packing

15 of 20

Predictive Model—Neural Network

  • The 28 prognostic variables are fed into a neural network with 4 hidden layers (16, 8, 8, 4) to predict patient outcomes from their H&E biopsies
  • 567 images were split into 80% training and 20% test sets
    • Overall Accuracy of 78.76%,
    • Sensitivity of 72.13% for favourable outcomes
    • Specificity of 86.54% for favourable outcomes

16 of 20

Predictive Model—Neural Network

  • A k-fold cross-validation was also performed with 10 folds to validate the model

  • Average overall accuracy: 77.72%

  • Average Sensitivity: 77.17%

  • Average Specificity: 78.12%

17 of 20

Advantages of the Workflow

  • Uses routine H&E stains:
    • No additional biomarkers
  • Low cost:
    • All aspects of the workflow use free, open-source software
    • No additional staining required
    • No whole-slide scanning required
  • Low computational cost:
    • Each patient was analysed in under 2hrs with a standard laptop with 8GB RAM
  • Time efficient:
    • Does not delay diagnosis and initiation of cancer treatment
    • Does not require pathologist to manually analyse slides for various prognostic factors
  • Quantitative:
    • Not subject to operator bias or semi-quantitative manual assessment

18 of 20

Applying the Workflow to TILs in Breast Cancer

  • A screenshot from one of the biopsies from The Cancer Genome Atlas, run through the workflow (https://mathbiol.github.io/tcgatil/)
  • The pipeline is fairly modular and it is possible to only segment and identify TILs, allowing it to be applied to other cancer types

19 of 20

Other Potential Applications

  • Single-cell segmentation and classification allows us to obtain spatial maps of the various cell types

20 of 20

THANK YOU!