Sample sheets & sequencing data QC
Setting up & QC a sequencing run
Create sample sheet
This will be the name of your sequencing run on BaseSpace.
For workflows that use dual indexing (Nextera, TruSeq Custom Amplicon, ..), this field is required (enter amplicon). For non-indexed/single-indexed TruSeq RNA or TruSeq DNA libraries, leave blank.
REQUIRED. This must match an official analysis workflow as set by Illumina. GenerateFASTQ tells it to just generate FASTQ files and not do more analyses.
# cycles to perform for read 1 (read length)
# cycles to perform for read 2, for paired-end runs (read length)
Optional fields that CZ Biohub uses internally, not required
Indexes or barcodes are usually 8 bp long, if you use custom longer indexes/barcodes you need to mention it here (if so, remember to also reduce the number of cycles for read 1 and 2 accordingly: e.g. for 12 bp barcodes, you need to reduce read1 and read2 cycles to 146 each (to account for the 4 extra cycles used per to detect each longer index).
[Data] section (Biohub)
Study_description: Description of a project to which the current sequencing run belongs. Special characters and commas are not allowed. (Required)
BioSample_ID: A short ID assigned to the specific biological sample based on the conventions used at your institution/group. Accepted characters include numbers,
letters, "-", and "_".
Biosample_Description: Description (longer than 10 characters) of the specific biological sample. The description should be the same for all the samples derived from this BioSample. Special characters and commas are not allowed. (Required)
Sample_ID: A short ID assigned to the specific library. Accepted characters include numbers, letters, "-", and "_". Sample_ID must begin with a letter. This field should be unique for each row. Sample_ID can be the same as Sample_Name. If so, fastq files will be demultiplexed into the same folder. Otherwise, fastq files belonging to the same library will be demultiplexed into a folder whose name is Sample_ID (Required)
Sample_Name: A distinct and descriptive name for each specific library. Accepted characters are numbers, letters, "-", and "_". Name must begin with a letter. This field should be unique for each row. Sample_Name will become the final fastq file names. (Required)
Sample_Owner: This column is used to keep track of to whom each library belongs to. Please fill in following the format FirstName_LastName. (Required)
Index_ID and Index2_ID: Name of the first or second index. Accepted characters are numbers, letters, "-", and "_". This field is left blank only when the number of libraries in pool is 1 and index read lengths are 0. (Required)
Index and Index2: Sequence of the first index. This field is left blank only when the number of libraries in pool is 1 and index read lengths are 0. (Required)
Organism: This field is used to record the organism the DNA comes from. Some examples are Human, Mouse, Mosquito, Yeast, Bacteria, etc. For individual species, please spell out the entire scientific name. Special characters are not permitted. (Required)
[Data] section lists sample names and index/barcode sequences.
For GenerateFASTQ workflows, the following 4 columns are required for MiSeq sample sheets (see guide here):
Make sure the sample sheet is saved as a .csv file!!
Create sample sheet using Illumina Experiment Manager
Important note: If you use index adapters or barcodes from a source other than Illumina, you have to make your own sample sheet by modifying an existing sample sheet, and cannot use Illumina Experiment Manager.
This is where we tell the machine that all we need is the FASTQ files with the read data, no other analyses required.
Create sample sheet using Illumina Experiment Manager
9. Create a New Sample Plate: click on ‘New Plate’
10. Under Sample Plate, click on ‘Select All’ and then ‘Add Selected Samples’
11. Check if all is correct !!
12. When finished, click ‘Finish’ to save the sample sheet.
Crucial step
Example MiSeq sample sheet generated through Illumina Experiment Manager
Automatically added based on Index Adapters kit that was selected
General tracking info about the run
This tells the machine which type of analysis to perform on the data
Library prep kit and Index adapters kit used
For GenerateFASTQ workflows, ‘Amplicon’ needs to be entered in this field
Setting up & QC a sequencing run
Setting up & QC a sequencing run
How to QC a sequencing run in BaseSpace
How to QC a sequencing run in BaseSpace
What is overclustering/underclustering?
A cluster is a clonal group of a library fragment generated in the first step of sequencing. Each cluster should ideally yield one (paired-end) read.
Overclustering = too much material was loaded on the flow cell, making that the resulting clusters are too close to each other. When overclustered, the sequencer cannot properly distinguish one cluster from another, leading to mixed signals and many reads failing the initial QC filter (low % PF). The high level of signal due to overclustering will also cause difficulties for the sequencer to properly focus. This high background signal will reduce base call quality (low %Q30).
Underclustering = a low amount of material was loaded on the flow cell. This reduces the number of clusters and therefore also reads the sequencing run will generate. Underclustering generally does not affect Q30.
Overclustering is worse than underclustering!
How to QC a sequencing run in BaseSpace
After logging into BaseSpace, Click on ‘Runs’ to view a list of all your runs.
Click on a Run Name to see it’s metrics.
How to QC a sequencing run in BaseSpace
How to QC a sequencing run in BaseSpace
How to QC a sequencing run in BaseSpace
Click on ‘Charts’ to view the QC charts
How to QC a sequencing run in BaseSpace
Flow Cell Chart: select % occupied under ‘Chart’
-> Signs of overclusterig or underclustering?
How to QC a sequencing run in BaseSpace
Data by cycle chart: select % Base under ‘Chart’ -> Adapter dimers? Short fragments?
Read 1 = 150 cycles
Read 2 = 150 cycles
Index ✕2
How to QC a sequencing run in BaseSpace
Data by cycle chart: select % Base under ‘Chart’ -> Adapter dimers? Short fragments?
Read 1 = 150 cycles
Read 2 = 150 cycles
Index ✕2
Variation in beginning of reads is normal
Signal around 25% (if 50%GC content)
Lower sequence diversity of indexes causes larger peaks
How to QC a sequencing run in BaseSpace
Average signal for GC and AT can differ based on GC content, but overall GCAT signal average should be 25%
How to QC a sequencing run in BaseSpace
Increase in G signal at end of reads
Issue 1 - short reads: Increase in G signal at end of reads
How to QC a sequencing run in BaseSpace
The first ~30 bases are the first portion of the illumina adapter, followed by heterogeneity from the barcode sequences (8-12 bases), followed by the last ~30 bases of the adapter. Then a large spike in the ‘dark’/no signal base ‘G’.
Issue 2 - adapter dimers: Large peaks for first 70 cycles, and then an increase in G signal
Adapter dimers being sequenced
How to QC a sequencing run in BaseSpace
Issue 3 - adapter dimers and short reads: Large peaks for first 70 cycles followed by more stable signal and then a steady increase in G signal
Adapter dimers being sequenced
How to QC a sequencing run in BaseSpace
On the ‘Metrics’ page (if using PhiX)
How to QC a sequencing run in BaseSpace
Scroll down to ‘Per Lane Metrics’
Number of paired end reads that were generated
How to QC a sequencing run in BaseSpace
On the ‘Indexing QC’ page:
How to QC a sequencing run in BaseSpace
...
...
...
...
If correctly pooled, all samples should have more or less the same # of reads assigned
How to QC a sequencing run in BaseSpace
Info on finding troubleshooting and finding the demux file in BaseSpace here
-How to edit sample sheet and requeue in Basespace here
Info on finding troubleshooting and finding the demux file using Local Run Manager here
Info on finding troubleshooting and finding the demux file using MiSeq Reporter here
-How to edit sample sheet and requeue using MiSeq Reporter here
Setting up & QC a sequencing run
How to QC a sequencing run in Sequence Analysis Viewer
Find Illumina Sequence Analysis Viewer Software Guide here
Same QC metric principles as described for BaseSpace apply!
How to QC a sequencing run in Sequence Analysis Viewer
After opening Sequence Analysis Viewer, click on ‘Browse’ and select the appropriate Run folder to view it’s metrics.
How to QC a sequencing run in Sequence Analysis Viewer
In the Flow Cell Chart, click on the arrow to display the ‘% Occupied’ data instead of ‘Intensity’.
Determine if the optimal amount of library was loaded onto the flow cell. Are there any signs of over/under clustering?
The % Occupied ranges here from 71 to 76. This indicates minor underclustering. This does not necessarily affect sequence quality, but it does reduce the maximum yield of data that can be obtained from a sequencing run. For the next sequencing run the researchers can load more material on the flow cell (provided that the same quantification protocol is being used) to generate more sequencing data from a single run .
For iSeq runs:
How to QC a sequencing run in Sequence Analysis Viewer
In the ‘Data by Cycle’ graph, select ‘% Base’ to determine if you had any adapter dimers, short fragments or potentially other issues.
How to QC a sequencing run in Sequence Analysis Viewer
Click on the ‘Summary’ tab to view how many reads were generated, and if the % aligned and the error rate are within the expected range.
% aligned should be similar to %PhiX that was spiked in
Error rate (%) should be <1.0
Cluster Count PF (M) shows how many paired-end passed filter reads were generated (in this case 5.03 million PF reads)
How to QC a sequencing run in Sequence Analysis Viewer
Click on the ‘Indexing’ tab to view metrics on demultiplexing.
% Undetermined reads = 100 – ‘%Reads Identified (PF)’
(= 9.07% in this example)
Check if all samples are equally represented (and none are missing)
Setting up & QC a sequencing run
SARS-CoV-2 only
for now
Check out the CZ ID help center for guides on how to QC and visualize consensus genomes
CZ ID will generate the consensus genomes and will allow you to view QC metrics in CZ ID itself! Follow the guide in the link above.
Alternatively, you can also download the consensus genome and the related QC files from CZ ID (see next slides).
Three major questions:
Coverage plots
Coverage-plots folder
Coverage depth: on average, how many reads cover each base in the genome?
Coverage breadth: how much of the genome has enough coverage depth to perform base calls?
Two gaps with 0x coverage
3 regions with <10x coverage, will show up as gaps in IDSEQ
No coverage at beginning and end of genome is normal
Coverage plot sample 1
Coverage plot sample 2
~200x depth on average
~35000 x depth on average
QUAST/*sample*/report.txt
Total length of the consensus genome should be close to the size of the SC2 genome = 29,903 bp
% of the genome that is covered (coverage breadth), should be >95%
Good coverage breadth!
call_consensus-stats/combined.stats.tsv
Sample a1 had 36,897x coverage (7,851,663 mapped reads)
Sample a2 had 238x coverage (52,383 mapped reads)
Setting up & QC a sequencing run
QC sequences and identify SNVs, gaps and lineage calls in Nextclade.
Nextclade - upload combined fasta file on https://clades.nextstrain.org/
Visualisation of identified mutations and gaps
Nextclade - upload combined fasta file on https://clades.nextstrain.org/
Missing bases
Ambiguous bases
Mutations vs
ref genome
Lineage call
QC results (see next slide)
(water control)
Nextclade: Phylogenetic-based sequence QC
Number of sites where a base could not be called: Areas with low or no sequencing coverage have no information to tell you which base should be at that site. These sites are labelled with N’s. When a sequence has two many N’s it is both hard to align and place on the tree, and thus they are removed from analyses. By default Nextstrain will drop sequencuences with less than 27,000 non ambiguous bases.
Mixed sites: If many sequencing reads support more than one base at a site, those sites will be designated with an IUPAC ambiguity code, that tells you which set of mutations were found at the site. While this can happen given a co-infection event, it more commonly occurs due to sample cross-contamination.
Private mutations: If a sequence differs from the Wuhan reference genome by (currently) more than 20 mutations, it will be flagged as having a high number of “private” mutations. The threshold for flagging a sequence as problematic will be changed as the diversity of SARS-CoV-2 increases over the pandemic.
Clusters of mutations: If your sequence has one or more areas with 6 mutations within a 100nt wide window, then that will be considered a “cluster of mutations” and it will be flagged unless it occurs at a recognized area of the genome. Such clusters of mutations are often artefactual, resulting from challenges aligning the sequence.
Troubleshooting too many N’s
N
Raw Reads Aligned to Reference
Reference Genome
Consensus Sequence
NNNN
Gap
pipeline requires >10 reads to call a base
Troubleshooting ‘mixed sites’
M
Raw Reads Aligned to Reference
Reference Genome
Consensus Sequence
C
M
pipeline requires a base to be >90% present to be called
Troubleshooting private mutations
Raw Reads Aligned to Reference
Reference Genome
Consensus Sequence
A
C
C
G
P
Troubleshooting clusters of mutations
C
Raw Reads Aligned to Reference
Reference Genome
Consensus Sequence
(100 bp)
A
C
C
G
A
G
T
C
G
A
A
C
Frameshift mutations
Other QC checks
Nextclade/Nextstrain
Clades: naming explained here: https://docs.nextstrain.org/en/latest/tutorials/SARS-CoV-2/steps/naming_clades.html
Sample a1 & a2 -->
How to requeue demux in BaseSpace
On the Summary page of the run, click on the hourglass, select Requeue and then select Sample Sheet
Click with your cursor in the black part of the screen and select everything (ctrl+A or cmd+A)
Click delete to delete the content of the sample sheet
Open your corrected sample sheet (saved as a .csv file) in a text editor such as Notepad or TextEdit, select all of its content (ctrl+A or cmd+A) and copy it into the black screen.
A green banner saying “Sample sheet is valid” will re-appear automatically if the copied sample sheet is valid. If not, make sure there are no columns or data fields missing in your sample sheet and that you indeed copied the full contents of the sample sheet into the field.
If sample sheet was valid, click on the blue button for Queue Analysis