Key concept: OTU or ASV
Do we really know what an OTU is or it should be?
What did the term OTU refer to before the era of DNA sequencing?
What is an OTU in DNA metabarcoding era?
DNA metabarcoding data analysis
Every time we analyse NGS data we need to face the issue related to sequencing errors.
For instance in genomics and transcriptomics we usually trim-out low quality region of reads.
In DNA-metabarcoding we need to account for:
DNA metabarcoding data analysis
OTU picking is the bioinformatic procedure applied for several years to treat DNA-metabarcoding data, with two general purposes:
Historically, OTUs represent groups of sequences with a similarity over a defined threshold, usually 97%.
We will briefly discuss how this threshold was defined and the reason why we need to overtake the idea of a fixed threshold.
What is a species?
A species is often defined as the largest group of organisms in which any two individuals of the appropriate gender or mating type can produce a fertile offspring, typically by sexual reproduction.
Ciccarelli FD et al (2006) Toward automatic reconstruction of a highly resolved tree of life. Science. 3;311(5765):1283-7.
Some definitions
OTU
OTU
Relationship between the rRNA 16% identity percentage and genomic DNA hybridization between different couples of microorganisms belonging to the Bacteria domain.
Orange area: 16S identity % and % of genomic similarity, higher than 97% and 70% respectively. Same species.
Blue area: 16S identity % higher than 97% but genomic similarity lower than 70%. Species sharing a similar 16S but not the same genetic content.
Green area: 16S identity % and % of genomic similarity, lower than 97% and 70% respectively.
TAKE HOME MESSAGE: The 16S gene may be very similar but does not always reflect the entire genome.
TAKE HOME MESSAGE: The 16S rRNA gene may be very similar but does not always reflect the entire genome.
Thus be careful using predictive software such as:
-1) PICRUSt2 for prediction of metagenome functions
-2) Tax4Fun2: prediction of habitat-specific functional profiles and functional redundancy based on 16S rRNA gene sequences
OTU (Operational Taxonomic Unit)
First application for OTU definition were based to agglomerative algorithms, a bottom-up method for obtaining hierarchical clustering.
Observed sequences
OTU (Operational Taxonomic Unit)
In the first step the agglomerative algorithm creates a cluster per sequence:
OTU (Operational Taxonomic Unit)
Then it starts to cluster sequence according to the imposed similarity threshold (i.e. 97%):
OTU (Operational Taxonomic Unit)
Then it tries to find the closest clusters and combining them. This operation is called linkage.
OTU (Operational Taxonomic Unit)
Then it tries to find the closest clusters and combining them.
OTU (Operational Taxonomic Unit)
The agglomerative algorithm requires a definition of linkage: how to calculate the distance between when a cluster contains more than one sequence:
OTU
The advent of NGS technologies and its usage in metabarcoding investigation led to the development of new approaches:
Due to read length limitation, we are just analyzing portion of the 16S.
OTU
Reference-based OTU picking. Sometimes referred as “consistent labelling” or “Phylotypes”
Reference Collection
Amplicon sequences
Global Alignment
OTU
Reference-based OTU picking
OTU
denovo OTU picking: this approach doesn’t require any reference collection. It builds OTUs just considering the similarity between sequences.
Amplicon sequences
Pairwise Alignments
OTUs
singletons
OTU
Regardless to the applied OTU picking procedure the main result is an OTU table:
Moreover, for each OTUs a representative sequence is defined
| sample 1 | sample 2 | sample 3 | sample 4 | sample 5 |
OTU 1 | 0 | 44 | 5 | 7 | 0 |
OTU 2 | 20 | 0 | 100 | 19 | 21 |
OTU 3 | 3 | 33 | 0 | 3 | 0 |
OTU 4 | 0 | 2 | 0 | 25 | 0 |
OTU 5 | 0 | 0 | 0 | 0 | 0 |
OTU
OTU1
OTU2
OTU3
OTU4
OTU5
OTU6
OTU
| PRO | CONS |
Reference-based |
|
|
denovo |
|
|
OTU - drawback
Examples: OTU-based surveys of human microbiome are unable to discern among dangerous and not dangerous members belonging to the Neisseria genus.
ASV
Amplicon Sequence Variants (ASV) is a recently developed approach allowing to infer the “real sequences” in the analysed samples, without imposing any similarity threshold.
Actually, it represents an evolution of denoising method developed for 454 data (e.g. AmpliconNoise PMID: 21276213) treatment in order to be applicable to Illumina sequences
ASV
The real advantage of ASV relies in the fact that it does not produce OTU!!!
ASV methods infers the real sequences in samples generating the reads we are analyzing. This real sequences are called features.
This mean that are not related to:
Features are comparable between different studies/surveys because they refer to real object.
ASV
DADA2 builds an error model to explain if the difference observed between sequences are due to real variability or are PCR/sequencing artifacts.
The error model is not pre-calculated but it is strictly defined by observing the real data. It is inferred by comparing the sequences.
While it is observing the sequences to infer the error model it identifies and removes chimeric sequences.
PAY ATTENTION: different sequencing run must be analysed in different DADA 2 runs. This is due to the fact that different runs generate different error profiles.
The result of the analysis is a “features table” a table (the same structure of OTU table) and the “feature sequences”
ASV
DENOISING
ASV
Error model
Sequences Abundance
ASV1
ASV2
ASV3
ASV5
ASV4
A
T
TT
A
A
T
T
G
T
AA
T
TT
T
TT
C
CA
ASV Table
| sample 1 | sample 2 | sample 3 | sample 4 | sample 5 |
ASV 1 | 0 | 54 | 6 | 7 | 0 |
ASV 2 | 20 | 0 | 96 | 19 | 21 |
ASV 3 | 3 | 33 | 0 | 3 | 0 |
ASV 4 | 0 | 2 | 0 | 24 | 0 |
ASV 5 | 0 | 0 | 0 | 0 | 0 |
Is Denoising suitable for all sequencing technologies?
Chimeras�
When you are dealing with metabarcoding data, probably the most important issue is the chimeras treatment.
Chimeras are amplicon generated from at least two different templates.
OTU
Chimera removal
Ref 2
Ref 1
Chimeric sequence
OTU
Chimera removal
Sequence 2
Sequence 1
Chimeric sequence
OTU - ASV comparison on a mock community�
We compared the two approach by using a mock community. The V5V6 hypervariable 16S regions were amplified and sequenced.
Species | % |
Bacteroides fragilis (ATCC 25285D-5) | 0.08 |
Phocaeicola vulgatus (ATCC 8482D-5) | 0.093 |
Bifidobacterium adolescentis (ATCC 15703D-5) | 0.067 |
Clostridioides difficile (ATCC 9689D-5) | 0.16 |
Enterococcus faecalis (ATCC 700802D-5) | 0.053 |
Lactiplantibacillus plantarum (ATCC BAA-793D-5) | 0.067 |
Enterobacter cloacae (ATCC 13047D-5) | 0.106 |
Escherichia coli (ATCC 700926D-5) | 0.093 |
Helicobacter pylori (ATCC 700392D-5) | 0.027 |
Salmonella enterica subsp. enterica (ATCC 9150D-5) | 0.093 |
Yersinia enterocolitica (ATCC 27729D-5) | 0.093 |
Fusobacterium nucleatum subsp. nucleatum (ATCC 25586D-5) | 0.067 |
OTU - ASV comparison on a mock community�
Data import
Data quality visualization
Denoising
ASV table
PE merging
OTU Picking
Chimera Checking
OTU table
97%
98%
99%
Primer trimming
OTU - ASV comparison on a mock community
| Input | Primer trimming | merging | Not Chimera |
ASV | 378,621 | 368,511 (92.67%) | 294,317 (79.92%) | 229,600 (62.3%) |
OTU 97% | 324,218 (87.98%) | 274,564 (72.52%) | ||
OTU 98% | 273,623 (72.27%) | |||
OTU 99% | 266,906 (70.49%) |
OTU - ASV comparison on a mock community
OTU - ASV comparison on a mock community
How to denoise with QIIME2
qiime dada2 denoise-paired \
--i-demultiplexed-seqs demux-paired-end.qza \
--p-trim-left-f 13 \
--p-trim-left-r 13 \
--p-trunc-len-f 150 \
--p-trunc-len-r 150 \
--p-n-threads 2 \
--o-table table_16S.qza \
--o-representative-sequences rep-seqs_16S.qza \
--o-denoising-stats denoising-stats_16S.qza \
--p-n-reads-learn 50000 \
--verbose \
--p-max-ee-f 2 \
--p-max-ee-r 2
Qiime2 implements dada2 to perform data denoising. 4 version are available:
OK… This is the theory, but what about real samples?
OTU or ASV?
The answer is:
«It depends on your project and the target organisms»
To ensure accurate results when using DNA metabarcoding, it's important to have a solid understanding of the biology of your target group beforehand.
Challenges in ITS-based Fungal Community Studies�
Kauserud (2023) ITS alchemy: On the use of ITS as a DNA marker in fungal ecology. Fungal Ecology ISSN 1754-5048.
Challenges in ITS-based Fungal Community Studies�
Kauserud (2023) ITS alchemy: On the use of ITS as a DNA marker in fungal ecology. Fungal Ecology ISSN 1754-5048.
DADA2 and ASVs�
�
Kauserud (2023) ITS alchemy: On the use of ITS as a DNA marker in fungal ecology. Fungal Ecology ISSN 1754-5048.
The Myth of exactness
�
Kauserud (2023) ITS alchemy: On the use of ITS as a DNA marker in fungal ecology. Fungal Ecology ISSN 1754-5048.
ASVs and OTUs Integration
�
Kauserud (2023) ITS alchemy: On the use of ITS as a DNA marker in fungal ecology. Fungal Ecology ISSN 1754-5048.
The dream is to use DADA2 and then cluster my ASVs with a specific cutoff (e.g. 97%).
Is this possible?
ASV clustering with qiime vsearch
qiime vsearch cluster-features-de-novo
Data import
Data quality visualization
Denoising
ASV table
OTU (97%) table
97% 98% 99%
Primer trimming
ASV clustering
How to cluster at 97% the ASV with QIIME2
qiime vsearch cluster-features-de-novo \
--i-table table_16S.qza \
--i-sequences rep-seqs_16S.qza \
--p-perc-identity 0.97 \
--o-clustered-table table-16S-dn-97.qza \
--o-clustered-sequences rep-seqs-16S-dn-97.qza