1 of 34

Clustering and Differential Expression

Introduction to Single Cell RNA-Seq (45)

Timothy Tickle

Brian Haas

2 of 34

Agenda (Clustering and Differential Expression)

  • Dimensionality Reduction
    • PCA
    • t-SNE
  • Differential Expression
    • SCDE
    • MAST

3 of 34

Making Sense of Variation

4 of 34

Identifying Relevant, “Highly Variable” Genes

5 of 34

Variable Genes in Seurat

Calculate mean expression.

Calculate disperstion (standard deviation).

Calculate z-score for dispersions within each bin.

Stratifies and controls from the relationship between the variability and mean expression.

Default�Standard�Deviation

6 of 34

Dimensionality Reduction

•Start with many measurements (high dimensional).

  • Want to reduce to few features (lower-dimensional space).

•One way is to extract features based on capturing groups of variance.

•Another could be to preferentially select some of the current features.

  • We have already done this.

•We need this to plot the cells in 2D (or ordinate them)

•In scRNA-Seq PC1 may be complexity or technical.

7 of 34

Dimensionality Reduction

8 of 34

PCA: Overview

•Eigenvectors of covariance matrix.

•Find orthogonal groups of variance.

•Given from most to least variance.

  • Components of variation.
  • Linear combinations explaining the variance.

9 of 34

PCA: an Interactive Example

10 of 34

PCA: in Practice

Things to be aware of-

•Data with different magnitudes will dominate.

  • Zero center and divided by SD.

•(Standardized).

•Can be affected by outliers.

•Data is often first filtered to remove noise.

11 of 34

PCs

Notice how lower PCs look more and more “spherical” - this loss of structure indicates that the variation captured by these PCs mostly reflects noise.

12 of 34

How Many Components Should We Use?

Elbow Plot (Scree Plot)

13 of 34

t-SNE: Collapsing the Visualization to 2D

14 of 34

t-SNE: Nonlinear Dimensionality Reduction

15 of 34

t-SNE: How it Works

16 of 34

Visualizing t-SNE

17 of 34

PCA and t-SNE Together

•Often t-SNE is performed on PCA components

  • Liberal number of components.
  • Removes mild signal (assumption of noise).
  • Faster, on less data but, hopefully the same signal.

18 of 34

Plotting Metadata on Ordinations

Metadata

Gene Expression

X

X

19 of 34

Caution When Interpreting t-SNE

Nonlinear�Optimized for local distanct

Big clusters can just mean more cells.

20 of 34

Learn More About t-SNE

•Awesome Blog on t-SNE parameterization

•Publication

•Nice YouTube Video

•Code

•Interactive Tensorflow

  • http://projector.tensorflow.org/

21 of 34

Defining Clusters Through Graphs

22 of 34

Local Moving Heuristic

23 of 34

Agenda (Clustering and Differential Expression)

  • Dimensionality Reduction
    • PCA
    • t-SNE
  • Differential Expression
    • SCDE
    • MAST

24 of 34

Differential Expression

25 of 34

Differential Expression Analysis

Soneson and Robinson, 2017

Many of the DE methods developed for bulk RNA-seq (e.g. edgeR, DE-seq) have serious limitations when applied to scRNA-seq data because of dropouts, so apply with caution!

26 of 34

Single Cell Differential Expression (SCDE)

27 of 34

Singe Cell Differential Expression (SCDE)

28 of 34

SCDE is Much More Sensitive and Specific

One of the disadvantages of SCDE is its run-time, which does not scale well for large datasets. Newer methods like MAST (Finak et al., 2016) overcome this!

29 of 34

MAST

•Uses hurdle model

  • Two part generalized linear model to address both rate of expression (prevalence) and expression.
  • GLM means covariates can be used to control for unwanted signal.

•CDR: Cellular detection rate

  • Cellular complexity
  • Values below a threshold are 0

Additionally introduces a�GSEA method

https://github.com/RGLab/MAST

30 of 34

MAST: Hurdle Models

31 of 34

Dot Plots

Size of circle

•Gene prevalence in cluster.

  • Color of circle

•More red, more expressed in cluster.

  • Scales well with many cells.

32 of 34

Seurat: Differential Expression

•Default if one cluster again many tests.

  • Can specify an ident.2 test between clusters.

•Adding speed by excluding tests.

  • Min.pct - controls for sparsity
  • Min percentage in a group
  • Thresh.test - must have this difference in averages.

33 of 34

Seurat: Many Choices of DE

Bimod�- Tests differences in mean and proportions.

Roc�- Uses AUC like definition of separation.

T�- Student's T-test.

Tobit�- Tobit regression on a smoothed data.

MAST�- Hurdle model for zero inflated data

….

34 of 34

Section Summary

We motivated dimensionality reduction with the helpfulness of focusing on higher variability.

We explored several methods for dimensionality reduction.

  • Contrasted and showed how to leverage together.

Explored differential expression.