1 of 13

ConsensusML for Cancer Biomarkers using Pediatric AML RNA-seq data - Final Hackathon Presentation

Team Members:

Jenny Smith (lead); Ronald Buie; Vikas Peddu; Ryan Shean; David Lee (technical); Sean Maden (scribe)

2 of 13

Methods Summary Cont. - ML algorithms summary

Algorithm Type

Resource Name

Support vector machines (SVM)

e1071

Random Forest, Gradient Boosting, XGBoost

xgboost; sklearn

Lasso

sklearn

glmboost

mlr

voomNSC, voomDLDA, PLDA, PLDA2, nblda, deepboost, blackboost, logitboost

MLSeq;DESeq2

3 of 13

Results: Normalized and Filtered Gene-level Expression

N = 1,998 genes (9.33% retained, adj. p-val <0.05)

Mean expr. increased (0.50 to 1.71)

Mean var. diff. increased (0.76 to 2.19)

4 of 13

Results: Selected Feature Consensus

Three models were built on the data, and top features extracted

The Lasso (N = 3212 loci), Random Forest (1147), and XGBoost (332)

N = 11 loci were deemed important by all three.

11

5 of 13

Results: Selected Feature Consensus, cont.

  • R packages MLseq and DEseq: Made specifically for logistic regression on RNAseq data
  • Models used:
    • voomNSC, voomDLDA, PLDA, PLDA2, and nblda
    • Deepboost, blackboost, logitboost were in the process of being tested
  • GATA2 and CSF3R were found by PLDA and PDLA2 (amongst 100 others)

6 of 13

7 of 13

Logistic Regression LASSO

Gene.name

Mod.RG

nonzero.coef

HOXA9

0.2184967372

PLCB4

0.0931645222

GOLGA8M

0.0868262987

ITGB3

0.0757604883

SLITRK5

0.0546755323

UNC13B

0.0280847887

KIAA1324L

0.0002750672

CROCC2

-0.0257766559

LINC01835

-0.0365784818

CLEC10A

-0.0430241193

GPR12

-0.0742916453

AC111000.4

-0.085920489

AC120498.2

-0.0956569772

8 of 13

Results: Selected Feature Consensus, cont.

GLMBoosted regression on Event Free Survival Time.

HPS4,HSPB1 - Heat shock proteins that have been implicated in relapse of HBV associated hepatocellular cancer

TPCN2 which has been associated with relapse of prostate cancer after surgery

9 of 13

Logistic Regression LASSO

  • 6% Test Error (N=44 AMLs )

True

Low risk

True

High Risk

Predicted

Low Risk

21

2

Predicted

High Risk

1

20

10 of 13

Support Vector Machine Analysis

  • SVM were used to test association between severity (risk) and genetic expression.
  • Linear SVM used 24 vectors.
  • Weighted predictions were used to score features for likelihood of high and low risk status.
  • Generally high scoring accuracy of 93% on test data
  • For simplicity of analysis, the 4 highest (2 associated with high risk, and 2 with low risk) performing features were reviewed in the literature.

11 of 13

3 out of 4 ain’t bad! (i.e. we should check the others)

AC004080.4

DTD1

FOLH1

ST18

12 of 13

Thanks to:

Ben Busby

Kate Hertweck

Hackathon presenters

Hackathon participants

Fred Hutch, AWS, and NCBI

13 of 13

Thanks to:

Ben Busby

Kate Hertweck

Hackathon presenters

Hackathon participants

Fred Hutch, AWS, and NCBI