1 of 61

Experimental design and instruments

2 of 61

Experimental design

  • A study is valuable may strongly rely on the experimental design.
  • Several factors need to be considered
    • Sample groups – how to define the control and experiment groups.
    • Sample types – what kind of samples need to be used, bacterial, tissue, blood, etc.
    • Techniques – sequencing strategy
  • A thorough discussion between biologists, physicians and bioinformaticians is very crucial.

3 of 61

Sample groups

  • Defining sample groups needs to match the research topic.
  • The key factor for comparison should cause a clear separation between control and experimental groups.
  • The other factors should be as diverse as possible, and the variance between groups should be small.
  • The common factors which need to be considered are age, gender, nation, health, etc.

4 of 61

Exercise I

  • We want to do a study for understanding the pathogenesis of adenocarcinoma, how should be define the groups.
    • Age
    • Gender
    • Nation
    • Health
    • Stage

5 of 61

Exercise II

  • A hypothesis is a specific gene may cause Alzheimer disease. We want to validate this hypothesis and also try to understand how the cell regulations are influenced by this gene. How should be setup the groups?
    • Age
    • Gender
    • Nation
    • Health
    • The duration of Alzheimer disease

6 of 61

Replicates

  • Replicates are the repeated experiments. There are several reasons for performing replicates
    • Avoid bias or special cases – a sample shows a special phenomena cannot represent the whole set. Ex: a student get A+ in a class, it does not all students get A+.
    • Reproducibility – if a phenomena does exist, it should be reproduceable. The phenomena only shows one time does not mean it do exist, ex: contaminant.
    • Statistical analysis – in order to make the results comparable, performing statistical analysis is required. Statistics need multiple samples for computing significance.

7 of 61

Biological replicate

  • Biological replicate is repeating the same experiment by using different samples from a group.
  • For example, selecting three colonies of E. coli which have a specific gene knockout. The colonies are different ones, but the experimental condition/group is the same.
  • Biological replicate is mainly for avoiding the issue of special case in order to understand the general situation of the experimental conditions.

8 of 61

Technical replicate

  • Technical replicate is performing the same experiment with identical samples.
  • For example, selecting a colony of E. coli which have a specific gene knockout. Using the same colony to perform the experiment multiple times.
  • The purpose of setting technical replicate is for minimizing the influences caused by technical issue. Or the samples which cannot have biological replicates.
  • In general, biological replicates are more important than technical replicates

9 of 61

10 of 61

Sample preparation

  • Sample preparation is the process of getting DNA or RNA ready for Next Generation Sequencing (NGS).
  • NGS is being asked to handle more challenging samples, from diverse origins, of lower quality or of small size. Before these samples can be analyzed, they must be treated and prepared.
  • This helps to prevent contamination, improve accuracy and minimize the risk of biases.
  • Different types of genetic material (DNA or RNA) and the species or strains of samples have slightly different sample preparation processes. On top of that, the different applications of NGS add another dimension. Therefore, no preparation protocol is always optimal.

11 of 61

Different Types of Sequencing

  • Whole genome sequencing
    • The process of determining the entire DNA sequence of an organism’s genome at a single time. Nearly any biological sample containing a full copy of the DNA can provide the genetic material necessary for full genome sequencing.
  • Whole exome sequencing
    • A technique for sequencing all of the protein-coding regions in the genome, known as the exome. The goal is to identify genetic variants that alter protein sequences and are responsible for diseases. The exons need to be selected before they can be sequenced – these are called target-enrichment strategies.

12 of 61

Different Types of Sequencing

  • Targeted sequencing
    • This allows the sequencing of specific areas of the genome for in-depth analysis more rapidly and cost effectively than whole genome sequencing. This method typically requires less sample input than other sequencing types.
  • RNA sequencing
    • Reveals the presence and quantity of RNA in a sample at a given time. This allows the analysis of the transcriptome – the set of all coding and non-coding RNA transcripts. This means that post-transcriptional modifications, gene fusions, mutations and changes in gene expression over time can be explored. During library preparation, RNA is reverse transcribed to complementary DNA (cDNA), because DNA is more stable, and this allows for amplification using DNA polymerase.

13 of 61

Sample preparation

Extraction of genetic materials

Library preparation

Amplification

Purification and quality control

14 of 61

Extract the genetic material

  • This is the first step in every sample preparation protocol. Nucleic acids (DNA or RNA) are extracted from a variety of biological samples.
  • The source of samples
    • Homogenous population of cells
    • Fresh samples. If not possible, proper storing is needed, such as freezing.
  • Different cell types and used machine may have different ways to extract the samples.

15 of 61

Typical steps for extracting nucleic acids

  • Cell disruption
    • The first step in nucleic acid isolation is to break apart the cell wall or cell membrane to release the genetic material.
    • Physical methods: use force and involve some sort of grinding or crushing. They are usually used on more structured materials, such as tissues.
    • Chemical methods: disrupt cellular membranes with a variety of agents that denature proteins.
    • Enzymatic methods: usually combines with chemical and physical procedures. Typical enzymatic treatments include lysozyme, proteinase K and lipase.

16 of 61

Typical steps for extracting nucleic acids

  • Removal of membrane lipids and proteins
    • This is usually done by centrifugation, filtration or bead-based methods.
  • Nucleic acid purification
    • The commonly used approaches
      • Silica
      • Ion exchange
      • Cellulose or precipitation-based methods.
    • Wash buffers usually contain alcohol and are used to eliminate contaminants from the sample.

17 of 61

Silica based purification

18 of 61

Ion exchange

19 of 61

Cellulose or precipitation-based methods

20 of 61

Library preparation

  • A sequencing library is essentially a pool of DNA fragments with adapters attached.
  • The typical steps of library preparation
    • cDNA library synthesis (for RNA library only)
    • Fragmentation
    • Attachment of primers, adapters and barcodes
    • Library quantification

21 of 61

Fragmentation

  • Fragmentation is the breaking of DNA strands into pieces, which can be done by using physical, chemical or enzymatic methods.
  • The commonly used methods
    • Acoustic shearing
    • Restriction endonucleases

22 of 61

Attachment of primers, adapters, barcodes

  • Primer are short oligonucleotide which can hybridizes with the sample DNA and defines the region that will be amplified.
  • DNA polymerase can add a nucleotide only onto a preexisting 3'-OH group, it needs a primer to which it can add the first nucleotide
  • Adapters are short, chemically-synthesized oligonucleotides that can be attached to the ends of DNA molecules.
  • The optimal ratio is 10 adapters to one fragment. Too many adapters cause adapter dimers to form. They will significantly lower the sequencing efficiency and quality.
  • Barcodes are also short oligonucleotides to identify where each nucleotide was originally located.

23 of 61

Design of primer

  • Length
    • Normally, 15-20 nts which can largely avoid the self-complementary, such as hairpin.
  • Melting temperature (Tm) = 4 °C x (G+C) + 2 °C x (A+T)
    • Around 50-60 °C.
    • When the Tm is too low, the primer may anneal incorrectly or not at all. A longer primer may help, but self-complementary needs to be considered as well.
    • A high Tm can be OK if there are not long strings (>3) of Gs or Cs that can bind quickly, often incorrectly, and very tightly.

24 of 61

Design of primer

  • GC content
    • The optimal GC content of a primer lies between 40 and 60 %, and primers should have two to three Gs and Cs at the 3' end (GC-lock) to bind more specifically to the template DNA.
  • The primer should not include poly base regions.
  • Four or more bases that compliment either direction of the primer should be avoided.

25 of 61

Attachment of primers, adapters, barcodes

26 of 61

Tagmentation

  • Tagmentation uses an engineered transposase enzyme to fragment the DNA and add specific adapters to both ends of the fragments, all at the same time.
  • However, this method is much more sensitive to the amount of DNA input compared to other fragmentation methods.

27 of 61

Library quantification

  • It provides the number of nucleic acids ready to be sequenced in a sample. This is key to obtaining high quality NGS data.
  • There are a variety of techniques that can be employed to determine the number of nucleic acids present in an NGS library, such as qPCR, Fluorometric quantification, .

28 of 61

The influence of fragmentation

  • In principle, the library should reflect the starting material as much as possible.
  • The shorter fragments are typically less specific in terms of alignment and so further decrease the complexity of a sample.
  • The longer fragments usually contain more mismatching nucleotides which may cause problems for read alignment.

29 of 61

RNA Sequencing Library

  • For generating RNA library, RNA needs to be isolated and converted into cDNA. This is so that the information can be input into an NGS platform.
  • DNA is more stable than RNA and it allows for amplification using DNA polymerases.
  • Poly A tail can be used to generate primer for constructing cDNA library.

30 of 61

RNA Sequencing Library

  • rRNA is the most abundant component of total RNA isolated from human cells and tissues – it comprises of up to 90% of an RNA sample. These must be removed from total RNA before sequencing to allow efficient gene detection.
  • Two main approaches
    • The selection of polyadenylated RNA (polyA) using oligo primers.
    • Targeted depletion of rRNA is particularly useful when studying transcripts that lack a polyA tail.

31 of 61

Amplification

  • This is an optional step, but it is usually required.
  • It is dependent on the application of NGS and the sample size.
  • Amplification becomes essential to obtain enough coverage for reliable sequencing for samples with small amounts of starting material.
  • Polymerase chain reaction (PCR) is a common method to increase the amount of DNA.

32 of 61

Purification and quality control

  • This step is usually necessary to remove any unwanted material that could hinder sequencing.
  • Some NGS platforms may have narrow size requirements, and so discarding too large or too small fragments can improve sequencing efficiency.
  • This ‘clean up’ is typically done by magnetic bead-based clean up or on agarose gels.
  • Quality control is the final process before proceeding to sequencing. Confirming the quality and quantity of DNA improves the confidence of sequencing data.

33 of 61

Purification and quality control

34 of 61

Purification and quality control

  • The quality of NGS is controlled by some factors
    • Missing value – for RNA-seq, the low expressed genes may not be expressed. This could be solved by increasing the read depth or doing imputation.
    • Error rate – often occurs at 3’ end.
    • Duplicate reads – if the sample size is small, using PCR to increase the amount of reads is required. But it will generate many duplicated reads which will decrease the quality of sequencing. This could be solved by bioinformatic tools or using specific PCR enzyme.

35 of 61

Error rate

Sequencing Accuracy = 99.9%

Sequencing depth = 10M

Sequencing length = 300nts

Nucleotides = 300nts x 10M = 3000M = 3x109

Error = 3x109 x 0.1% = 3x109 x 10-3 = 3x106 = 3000000nts

36 of 61

Error rate

A T G C G G T A C G A T A C G A T

A

T

G

C

G

G

T

A T G C G G T A C G A T A C G A T

A

T

G

C

G

G

T

T

C

G

A

T

A

A

C

G

T

A T G C G G T A C G A T A C G A T

A

T

G

C

G

G

T

C

G

A

T

A

A

C

G

T

A T G C G G T A C G A T A C G A T

A

T

G

C

G

G

T

C

G

A

T

A

A

C

G

T

A

A

37 of 61

FastQC

  • A widely used software for checking the quality of sequencing
  • Q = -10 log10(e)
    • e = error rate
    • If e = 0.1%, Q = 30

38 of 61

Per base sequence quality

Quality is too low

39 of 61

Sequence quality scores

Quality is too low

40 of 61

Sequence content

It is okay for RNA-Seq. the first 10-15 nts

is influenced by library kit. But not for

DNA-Seq

A%=T%, G%=C%

Roughly, the same ratio

41 of 61

Sequence GC content

Should be normal distribution

42 of 61

Per base N content

A specific position containing many unknown nts

43 of 61

Sequence duplication levels

Most of the reads only show one time

Some reads shows many times

For RNA-Seq, it is fine because transcripts

are reflect the amount of expressed RNAs

44 of 61

OVERREPRESENTED SEQUENCES

45 of 61

Sequence length distribution

Most of the read lengths are identical

Some short reads exist. If the result is from

the reads which removed adapters, it is fine.

46 of 61

Adapter content

No adapter

Containing adapters

47 of 61

Per tile content

Blue means the quality is at or above the average for that base in the run, hotter colors represent

the low quality. This may be due to the transient problems such as bubbles going through the

flowcell, or they could be more permanent problems such as smudges on the flowcell or debris

inside the flowcell lane.

48 of 61

49 of 61

Batch effect

  • If multiple sequencing runs are performed for a sample set, batch effect will occur.
  • For solving this issue, using spike-in or biostatistical normalization is a common way.
  • It is more commonly used for RNA-Seq data.

50 of 61

Spike-in

  • Spike-in controls are synthetic nucleic-acid sequences that are added to a user’s sample and constitute internal standards for subsequent steps in the next generation sequencing workflow.
  • The spike-in controls should not be able to mapped to the reference genome for avoiding the misinterpretation. Thus, the spike-in reads are normally from artificial gene or completely far relationship species, ex: E. coli vs human.
  • The amount and contend of spike-in controls should be exactly the same for all the experimental runs.

51 of 61

Spike-in

Interested human gene

Spike-in Drosophila gene

Experiment

4592

200

Exp 2

2010

88

Ratio

2.27

After normalization

Before normalization

Interested human gene

Spike-in Drosophila gene

Experiment

4592 / 2.27 = 2014

200 / 2.27 = 88

Exp 2

2010

88

52 of 61

Read depth and coverage

  • Sequencing depth
    • Total number of usable reads from the sequencing machine (usually used in the unit “number of reads” (in millions). Especially used for RNA-seq.
  • Coverage
    • Redundancy: number of reads that align to, or "cover," a known reference. It describes how often, in average, a reference sequence is covered by bases from the reads. This is an important information because multiple observations per base are needed to obtain to a reliable call.
    • Percentage coverage of a reference by reads. E.g. if 90% of a reference is covered by reads (and 10% not) it is a 90% coverage.

53 of 61

Read depth and coverage

54 of 61

Read length

  • Read length describes the average length of the sequencing reads produced.
  • For whole genome sequencing or species identification, longer reads are preferable. Longer read lengths are also essential for capturing insertions and deletions or for sequencing regions with a lot of redundancy.
  • Short reads are effective for applications aimed at counting the abundance of specific sequences, identifying variants within otherwise well-conserved sequences, or for profiling the expression of particular transcripts.
  • There is typically a trade-off between sequencing depth and sequencing length. Sequencing platforms that perform longer reads will typically provide less coverage, a higher error rate, and higher cost per base relative to short read sequencing.

55 of 61

Single and paired-end sequencing

  • Single-end reading
    • the sequencer reads a fragment from only one end to the other, generating the sequence of base pairs.
  • Paired-end reading
    • it starts at one read, finishes this direction at the specified read length, and then starts another round of reading from the opposite end of the fragment.
  • Paired-end reading improves the ability to identify the relative positions of various reads in the genome, making it much more effective than single-end reading in resolving structural rearrangements such as gene insertions, deletions, or inversions. It can also improve the assembly of repetitive regions.
  • Paired-end reads are more expensive and time-consuming to perform than single-end reads.

56 of 61

Paired-end VS single end

57 of 61

Single end

Paired end

Cost

Time

Information

Counts analysis

De nove seq

Splice variants

SNP

Cheaper

Faster

More

Good

Not worth

Better

Better

Better

58 of 61

First generation sequencing

  • First generation sequencing is developed by Frederick Sanger and coworkers in 1977. The main method is chain-termination method.
  • The Sanger method, in mass production form, is the technology which produced the first human genome in 2001, ushering in the age of genomics.

59 of 61

First generation sequencing

60 of 61

61 of 61

Next generation sequencing platform

https://www.illumina.com/content/dam/illumina-marketing/documents/applications/ngs-library-prep/for-all-you-seq-rna.pdf