1 of 47

Welcome to genPALS workshop 2021

17 September 2021

2 of 47

Time

Topic

Presenter

11AM – 11:30 AM

Welcome and introduction

Single-cell genomics overview

Sam Morabito

11:30 AM – 12:00 PM

Experimental overview of scRNA-seq

Liz Rebboah

12PM – 1PM

Lunch

Mendocino Farms

1PM – 2PM

Pre-processing raw sequencing data

Fairlie Reese

3 of 47

What is GenPALS?

Bioinformatic methods

Your data

You

genPALS

4 of 47

How am I supposed to figure out how to analyze my data???

5 of 47

How am I supposed to figure out how to analyze my data???

6 of 47

What will I learn by attending genPALS throughout Fall 2021?

7 of 47

Your first single-cell genomics analysis

a story in pictures

8 of 47

9 of 47

10 of 47

11 of 47

12 of 47

13 of 47

+

=

14 of 47

+

=

15 of 47

16 of 47

single-cell genomics starter pack

17 of 47

18 of 47

19 of 47

20 of 47

21 of 47

22 of 47

23 of 47

Why should I listen to you? I can just copy and paste from a tutorial, right?

24 of 47

25 of 47

26 of 47

snRNA-seq analysis of the prefrontal cortex of 32 individuals

27 of 47

A real single-cell dataset vs the PBMC tutorial dataset

  • Frozen postmortem tissue

  • Multiple individuals

  • Many different sequencing batches

  • Tissue from different sources

  • Variable quality of samples

  • Possibly many more cells, cell-types, cell-states, etc…

28 of 47

Re-processing the Zhou et al dataset using the scanpy tutorial code

29 of 47

Inspect quality of all 32 samples

30 of 47

Data normalization and scaling

scanpy.pp.log1p(data)

scanpy.pp.scale(data)

31 of 47

The curse of dimensionality

32 of 47

Feature Selection & Principal Component Analysis (PCA)

scanpy.pp.highly_variable_genes(data) scanpy.tl.pca(data)

33 of 47

Feature Selection & Principal Component Analysis (PCA)

33k genes

50 Principal Components (PCs)

Highly variable genes

34 of 47

Visualizing the top Principal Components

33k genes

50 Principal Components (PCs)

Highly variable genes

35 of 47

Non-linear dimensionality reduction (UMAP, tSNE, etc)

scanpy.pp.neighbors(data, n_pcs=30) scanpy.tl.umap(data)

36 of 47

UMAP colored by AD Diagnosis of patient samples

37 of 47

UMAP colored by Sample ID

38 of 47

UMAP colored by cell-type marker genes

39 of 47

Grouping cells into clusters

scanpy.tl.leiden(data)

40 of 47

UMAP colored by cluster assignment

41 of 47

UMAP colored by sequencing batch

42 of 47

UMAP colored by log(total UMI counts)

43 of 47

UMAP colored by percent mitochondrial reads

44 of 47

+

=

45 of 47

I don’t know what’s wrong with my data…

batch effects?

ambient RNA?

how many clusters???

doublets?

mitochondrial clusters?

UMAP or t-SNE?

what should I filter?

compare to published data?

cross-species analysis?

when do I stop analyzing?

46 of 47

Topics for Fall 2021 genPALS series

  • Pre-processing sequencing data

  • HPC, virtual environments, and more

  • Quality control

  • Normalization

  • Batch correction

  • Feature selection & dimensionality reduction

  • Data visualization

  • Differential expression analysis

  • Advanced topics (trajectory analysis, RNA velocity, etc)

47 of 47

Join our Slack workspace!

And go to the workshops_2021 channel!