1 of 15

Compositional Microbiome Data �and �Differential Abundance

2 of 15

The Problem With Rarefying Sequencing Data

3 of 15

4 of 15

https://www.frontiersin.org/articles/10.3389/fmicb.2017.02224/full?report=reader

5 of 15

LOSING COUNT INFO

6 of 15

TRUE ABUNDANCE

RELATIVE ABUNDANCE

SPURIOUS CORRELATIONS

7 of 15

Centred Log-Ratio (CLR) Transformation

  • DD methods largely discriminate between samples based on the most relatively abundant features in the data, not necessarily the most variable

  • This problem is exacerbated by sensitivity of methods to total read depth of a sample

  • Remember rarefying to control for read depth doesn’t fix this because all we do is lose precision about relative abundance

8 of 15

Centred Log-Ratio (CLR) Transformation

  • The starting point for any compositional analyses is a ratio transformation of the data.

  • Ratio transformations capture the relationships between the features in the dataset and these ratios are the same whether the data are counts or proportions.

  • Taking the logarithm of these ratios, thus log-ratios, makes the data symmetric and linearly related, and places the data in a log-ratio coordinate space

  • Thus, we can obtain information about the log-ratio abundances of features relative to other features in the dataset, and this information is directly relatable to the environment.

9 of 15

CLR Transformation

  • Thus, we can obtain information about the log-ratio abundances of features relative to other features in the dataset, and this information is directly relatable to the environment.

  • We cannot get information about the absolute abundances since this information is lost during the sequencing process

10 of 15

Benefits of CLR-Transformation

  • The clr- transformed values are scale-invariant; that is the same ratio is expected to be obtained in a sample with few read counts or an identical sample with many read counts, only the precision of the clr estimate is affected.

11 of 15

Issues with CLR-Transformation

  • CLR-transformation uses the Geometric Mean of abundances of features in a sample x

  • The G(x) cannot be determined for sparse data without deleting, replacing or estimating the 0 count values.

  • Fortunately, there are acceptable methods of dealing with 0 count values as both point estimates using zCompositions R package (Palarea-Albaladejo and Martín-Fernández, 2015), and as a probability distribution using ALDEx2 available on Bioconductor.

12 of 15

13 of 15

How Do we Quantify Absolute Abundance?

14 of 15

POSITIVE CONTROL SPIKE INS

15 of 15

qPCR

https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0227285