1 of 59

Comparative genomics

2 of 59

What is comparative genomics

  • Comparative genomics is a field of biological research in which the genomic features of different organisms are compared.
  • The genomic features may include the DNA sequence, genes, gene order, regulatory sequences, and other genomic structural landmarks.

3 of 59

The first comparative genomics study

  • In 1986, the first comparative genomic study at a larger scale was published, comparing the genomes of varicella-zoster virus and Epstein-Barr virus that contained more than 100 genes each.
  • The first high-resolution whole genome comparison system of microbial genomes of 10-15kbp was developed in 1998.
  • Saccharomyces cerevisiae, the baker's yeast, was the first eukaryote to have its complete genome sequence published in 1996.

4 of 59

How Are Genomes Compared?

  • A simple comparison of the general features of genomes such as genome size, number of genes, and chromosome number presents an entry point into comparative genomic analysis.
  • Finer-resolution comparisons are possible by direct DNA sequence comparisons between species.

How similar between mouse

and human in genome level

5 of 59

How Are Genomes Compared?

  • Comparison of discrete segments of genomes is also possible by aligning homologous DNA from different species.

6 of 59

Group definition and comparison

  • Comparison means multiple groups are needed.
    • Species – Evolution
    • Types of individuals – Disease
    • Cell types
    • Sample sources
    • Subtype of a biological sample group
  • How to design the experimental groups is the crucial factor for comparative genomics.

7 of 59

What is mutation

  • The most common applications of comparative genomics is mutation detection.
  • Wild type
    • An individual having the normal phenotype; that is, the phenotype generally found in a natural population of organisms.
  • Mutant
    • An individual having a phenotype that differs from the normal phenotype.
  • Allele
    • An allele is a variation of the same sequence of nucleotides that encodes the synthesis of a gene product at the same place on a long DNA molecule.

8 of 59

Allele comparison

  • Homozygous
    • Having two identical alleles for a particular trait.
  • Heterozygous
    • Having two different alleles for a particular trait.
  • Dominant allele
    • In a heterozygous condition, the allele that is expressed.
  • Recessive allele
    • In a heterozygous condition, the allele that is not expressed.

9 of 59

Dominant and recessive allele

10 of 59

What is variation

  • Variation refers to an individual that possesses characteristics different from the others of the same kind.
  • Genetic variation usually arises as a mutation in a gene that encodes a protein or an RNA.
  • Variations may be reflected in the phenotype of the individual, e.g. a difference in size, color or pattern, or may be detectable only by DNA or protein sequencing.
  • A given variant may be the result of changes in a single gene, or the consequence of interactions among multiple genes.
  • Populations of the same species living in different parts of the world may be characterized by specific sets of variations.

11 of 59

12 of 59

The situations of gene change

  • Mismatch
    • A nucleic acid/nucleic acids are replaced by another one/ones.
  • Insertion
    • Extra nucleic acid/nucleic acids are placed into the genome.
  • Deletion
    • A nucleic acid/nucleic acids are removed from the genome.
  • If the mutation is caused by a single point mutation, we call it SNP (single nucleotide polymorphisms).

13 of 59

Types of mutations – consequences

  • Silent mutations
    • Change a codon into a mutant codon that specifies exactly the same amino acid.
    • The majority of silent mutations change the third nucleotide of a codon.

14 of 59

Types of mutations – consequences

  • Missense mutations
    • Change a codon into a mutant codon that specifies a different amino acid.
    • Two types of missense mutation
      • Conservative missense mutation: substituted amino acid has chemical properties similar to the one it replaces, and then it may have little or no effect on protein function
      • Non-conservative missense mutations that cause substitution of an amino acid with very different properties and have more noticeable consequences.

15 of 59

16 of 59

Types of mutations – consequences

  • Nonsense mutations
    • Change an amino acid specifying codon to a premature stop codon.
    • The proteins are smaller than those encoded by wild-type alleles of the same gene. The shorter, truncated proteins lack all amino acids between the amino acid encoded by the mutant codon and the C terminus of the normal polypeptide.
    • As a result the mutant polypeptide will be unable to function if it requires the missing amino acids for its activity.

17 of 59

Types of mutations – consequences

  • Frame shift mutation
  • Insertion or deletion of nucleotides within the coding sequence.
  • If the number of extra or missing nucleotides is not divisible by 3, the insertion or deletion will skew the reading frame downstream of the mutation.
  • As a result, frame shift mutations cause unrelated amino acids to appear in place of amino acids critical to protein function

18 of 59

19 of 59

Types of mutations – consequences

  • Mutations outside the coding sequence
    • Occurred in outside of coding
    • Changes in any of these critical signals can disrupt the process. For example, promoters and termination signals in the DNA of a gene instruct RNA polymerase where to start and stop transcription.

20 of 59

Mutation detection by comparative genomics

  • Reference genome sequence VS DNA reads
    • The potential mutations can be detected.
  • Control VS Experiment
    • The mutated genes exist in experimental group but not in control can be found.

21 of 59

The way to detect mutation

ATAGACAGAACGATAAAGACGCTACT

Reference genome

ATAGAGAGAACGATAAAGACGCTACT

ATAGAGAGAACGATAAAGACGCTACT

ATAGAGAGAACGATAAAGACGCTACT

ATAGAGAGAACGATAAAGACGCTACT

ATAGACAGAACGATAAAGACGCTACT

ATAGAGAGAACGATAGAGACGCTACT

ATAGAGAGAACGATA TAGACGCTACT

ATAGAGAGAACGATAAAGACGCTACT

DNA reads

22 of 59

The way to detect mutation

ATAGACAGAACGATAAAGACGCTACT

Reference genome

ATAGA - AGAACGATAAAGACGCTACT

ATAGA - AGAACGATAAAGACGCTACT

ATAGA - AGAACGATAAAGACGCTACT

ATAGA - AGAACGATAAAGACGCTACT

ATAGACAGAACGATAAAGACGCTACT

ATAGA - AGAACGATAGAGACGCTACT

ATAGA - AGAACGATA TAGACGCTACT

ATAGA - AGAACGATAAAGACGCTACT

DNA reads

23 of 59

The way to detect mutation

ATAGACAGAACGATAAAGACGCTACT

Reference genome

ATAGACTAGAACGATAAAGACGCTACT

ATAGACTAGAACGATAAAGACGCTACT

ATAGACTAGAACGATAAAGACGCTACT

ATAGACTAGAACGATAAAGACGCTACT

ATAGACAGAACGATAAAGACGCTACT

ATAGACTAGAACGATAGAGACGCTACT

ATAGACTAGAACGATA TAGACGCTACT

ATAGACTAGAACGATAAAGACGCTACT

DNA reads

24 of 59

The way to detect mutation

25 of 59

The way to detect mutation

26 of 59

Comparison between two groups

27 of 59

Clinical applications

  • In clinical applications, comparative genomics is often used for detecting SNPs which occur in specific diseases. For the genes having such SNPs are called molecular biomarkers or biomarkers.

No SNP

A SNP

28 of 59

Example – HER2 mutation

29 of 59

Tumor mutation burden (TMB)

  • TMB is a genetic characteristic of tumorous tissue that can be informative to cancer research and treatment. It is defined as the number of non-inherited mutations per million bases (Mb) of investigated genomic sequence, and its measurement has been enabled by NGS.

Somatic mutation x 106

Total exonic bases with sufficient coverage

TMB =

30 of 59

The threshold of TMB

  • Evaluation of TMB:
    • Low TMB as ≤ 5
    • Intermediate TMB > 5 and < 20
    • High TMB as ≥ 20 and < 50
    • Very high TMB as ≥ 50

Patient A has 100 mutations/1MB

Normally human contains 233,785 exons, and 8.8 bases per exon in average.

TMB = 100 x 106 / 2,057,308 = 4.8607

31 of 59

Waterfall plot

TMB

32 of 59

Manhattan plot

  • List all SNP positions detected by WGS, Moreover using p-values to reveal the relationship between the phenotypes or diseases to the data.
  • All the SNPs can produce a p-value to represent how significant of such SNP influence the data.
  • Normally, we define the cutoff of p-value is –log10(5x10-8) = 7.3
  • Using –log10(0.05) = 1.3 will produce too many false positives due to the large data points.

33 of 59

Manhattan plot

34 of 59

Quantile-Quantile plot (QQ plot)

  • By Manhattan plot, we cannot distinguish the SNPs real connect to the disease or they are occur randomly.
  • In order to check how reliable for the detected SNPs, QQ plot was required.
  • QQ plot compare the observed p-values with the expected p-values to discover the reliable SNPs cutoffs.

35 of 59

Quantile-Quantile plot (QQ plot)

36 of 59

Comparative transcriptomics

  • Comparative transcriptomics is often used for searching the differential expressed genes which can be the factors for causing diseases.
  • For evaluating the differences between experimental conditions, statistical tests are required for computing FDR or q-values.
  • Moreover, normalization is also an important step.
  • It can also be used to check the expression patterns cross different species.

37 of 59

Quantification and identification

  • Quantification is for generating values to represent the expression of the genes or proteins
    • RNA-Seq – read counts
    • Proteomics – intensity
  • Identification is the genes or proteins which can be found by sequencing or mass spectrometry

38 of 59

Normalization

39 of 59

Normalization

40 of 59

Normalization

41 of 59

RPKM

  • If the length of a gene is longer, more reads can map on it.

42 of 59

Differential expression analysis

  • Differential expression analysis means taking the normalized read count data and performing statistical analysis to discover quantitative changes in expression levels between experimental groups.
  • It is the most widely used analysis method for comparative transcriptomics.
  • It is also often used for searching the biomarkers.

43 of 59

HER2 receptor

HER2 receptor

44 of 59

When are two values different

45 of 59

Replicates increase reliability

  • More replicates, more reliable mean estimates

  • Define a fold change threshold ? Is it significant?

46 of 59

Fold change not informative enough

  • The 3 proteins have the same fold change

  • Significance of protein B > signif. of protein A > signif. of protein C

47 of 59

T-test

  • Consider both variance between groups and variance within groups
  • Test-statistic:
  • p-value: how likely is it to get a certain value of t by chance
  • significant, if p-value <= significance level alpha, alpha = 0.01

 

48 of 59

P-value

49 of 59

Multiple hypothesis problem

  • Multiplicity problem:
    • 10,000 hypotheses are tested simultaneously
    • Increased chance of false positives

  • Suppose
    • None of the 10,000 proteins is differentially expressed
    • We expect 10,000 * 0.01 = 100 with p-value < 0.01

Individual pvalues are no significant findings

50 of 59

Benjamini-Hochberg correction

  • Compute p-values of all tests (m=10)
  • Order them in increasing order

Test if:

50

 

 

Rank i

P-value

Adjusted significance level

? Reject H0

1

0.00002

0.001

yes

2

0.00013

0.002

yes

3

0.00034

0.003

yes

4

0.00087

0.004

yes

5

0.001

0.005

yes

6

0.0054

0.006

yes

7

0.009

0.007

no

8

0.01

0.008

no

9

0.023

0.009

no

10

0.4

0.01

no

Significance Level:

Adjusted Significance Level

51 of 59

Permutation-based FDR

Compute test statistic & p-value

Permute data

Repeat many times

Control

Stimulus

Control

Stimulus

Compute test statistic & p-value

Protein A

Protein B

Protein C

Protein D

Protein E

MEASURED

PERMUTED

  • False discovery rate estimated by counting hits on permuted data

Careful with technical replicates !!!

52 of 59

Volcano plot

53 of 59

54 of 59

Combine the information from DNA-seq and RNA-seq

  • Expression quantitative trait loci (eQTLs) are genomic loci that explain variation in expression levels of mRNAs.
    • DNA-seq – detect the variants from the same position or the identical associated genes.
    • RNA-seq – compare the expression values between different variants
  • The influences of mutations can be measured.

55 of 59

cis and trans eQTLs

  • Based on the position of eQTLs with associated genes, eQTLs can be divided to two types:
    • cis-eQTLs – they are located near the gene-of-origin (gene which produces the transcript or protein), and also called local eQTLs.
    • trans-eQTLs – they are located distant from their gene of origin, often on different chromosomes, and also called distant eQTLs.
  • One eQTL can regulate more than one gene.

56 of 59

Different types of eQTLs

57 of 59

Different types of eQTLs

58 of 59

eQTL analysis for different variants

59 of 59

Manhattan plot + eQTL box plot