1 of 28

Multivariate Analyses of Microbial Communities

2 of 28

Goals of Next Two Days

  • Explore the huge variety of statistical tests available for the interrogation of microbial community data

  • Gain experience of choosing and using these tests within the R computing language

  • Be aware of the limits and pitfalls of using these tests on particular microbial datasets

3 of 28

Microbial Communities Analysis Workflow

Raw

Sequences

Quality

Trimming

Dereplicate &

Denoise

Chimera

Detection

Taxonomy

Assignment

1

4

3

2

5

BIOINFORMATICS

DOWNSTREAM ANALYSIS

4 of 28

Quality

Filtering

6

Rarefied

Libraries

7

Alpha

Diversity

8

Unrarefied

Libraries

10

Microbial Communities Analysis Workflow

BIOINFORMATICS

1-5

Beta

Diversity

9

Differential

Abundance

Analysis

11

T tests, GLMs,

GLMMS

Heatmaps, PERMANOVA, Ordinations,

Divergence, Indicator Analysis

DESeq2

ALDEX2

CLR-

Transformations

5 of 28

Microbial Communities Are Complex

6 of 28

Essential Reading

7 of 28

Part 1: Alpha Diversity Analysis

8 of 28

What Are We Measuring?

9 of 28

What Are We Measuring?

Alpha diversity metrics are CORRELATED

10 of 28

YUNGAY

BAQUEDANO

2 TRANSECTS

MULTIPLE SITES

PER TRANSECT

YU1

YU2

YUx

BA1

BA2

BAx

MULTIPLE SAMPLES

PER SITE

Sample 1

Sample 2

Sample x

Sample 1

Sample 2

Sample x

Our data are

STRUCTURED

Samples are

not independent

data points

11 of 28

Part 2: Beta Diversity Analysis

12 of 28

Beta Diversity

We want to understand the differences in microbial community composition among samples and among groups of samples

13 of 28

Beta Diversity

We want to understand the differences in microbial community composition among samples and among groups of samples

But how do we summarise microbial communities with potentially hundreds of traits (species) ?

14 of 28

Beta Diversity

To do this, we need

  • a DISTANCE metric (how dissimilar are samples)

  • an ORDINATION technique (how do we display those dissimilarities in 2 (ish) dimensions

15 of 28

Ordination

Ordination methods like PCA:

  1. Calculate new synthetic variables (principal components), which are linear combinations of the original variables that…
  2. account for as much of the variance of the original data as possible

16 of 28

Ordination Techniques

EXPLORATORY ORDINATIONS:

  • Principal Coordinates Analysis (PCoA): Like a Principal Component Analysis (PCA) but can use any dissimilarity matrix (Bray-Curtis, UNIFRAC) etc. Can’t handle predictors.

  • Non-Metric Multidimensional Scaling (NMDS): Unique ordination technique in which you choose number of axes to ordinate to. Rank based approach, not absolute distances. Can’t handle predictors.

INTERPRETIVE ORDINATIONS

  • Constrained Correspondence Analysis (CCA): Calculates axes that ‘constrain’ variation in distance matrix to that explained by candidate variables, and axes that ordinate variance not explained by those factors.

17 of 28

Ordination Metrics (Can Be) Correlated

18 of 28

TESTS of BETA DIVERSITY

Paliy & Shankar

19 of 28

TESTS OF BETA DIVERSITY

PERMANOVA

The null hypothesis tested by PERMANOVA is that, under the assumption of exchangeability of the sample units among the groups, H0: “the centroids of the groups, as defined in the space of the chosen resemblance measure, are equivalent for all groups.” Thus, if H0 were true, any observed differences among the centroids in a given set of data will be similar in size to what would be obtained under random allocation of individual sample units to the groups (i.e., under permutation). 

ANOSIM

The null hypothesis for the ANOSIM test is closely related to this, namely H0: “the average of the ranks of within‐group distances is greater than or equal to the average of the ranks of between‐group distances,” 

20 of 28

PERMUTATIONAL TESTING

AXIS 1

AXIS 2

Vegetation

No Vegetation

REAL DATA

Between

Group

Within

Group

21 of 28

PERMUTATIONAL TESTING

AXIS 1

AXIS 2

Vegetation

No Vegetation

REAL DATA

Between

Group

Within

Group

AXIS 1

AXIS 2

PERMUTED DATA

Between

Group

Within

Group

22 of 28

TESTS OF BETA DIVERSITY

  • Overall, ANOSIM and the Mantel test were very sensitive to heterogeneity in dispersions, with ANOSIM generally being more sensitive than the Mantel test.

  • In contrast, PERMANOVA ... [was] largely unaffected by heterogeneity for balanced designs.

  • For unbalanced designs, however, all of the tests were (1) too liberal when the smaller group had greater dispersion and (2) overly conservative when the larger group had greater dispersion, especially ANOSIM and the Mantel test.

  • For simulations based on real ecological data sets, PERMANOVA was generally, but not always, more powerful than the others to detect changes in community structure

https://esajournals.onlinelibrary.wiley.com/doi/10.1890/12-2010.1

23 of 28

The Importance of Metadata

https://royalsocietypublishing.org/doi/abs/10.1098/rspb.2021.0552

24 of 28

The Importance of Metadata

25 of 28

The Importance of Metadata

26 of 28

The Importance of Metadata��Statistical Power and Data Completeness�

EXPERIMENTAL DESIGN

27 of 28

The Importance of Metadata��Statistical Power and Data Completeness�

DATA COLLECTION

28 of 28

The Importance of Metadata��Statistical Power and Data Completeness�

LAB ISSUES