1 of 72

CNV/SV Detection and Exome data analysis

25.03.2025

German Demidov

2 of 72

3 of 72

Remark about the expected level of participants

3

4 of 72

“how it came about that you are working in Rare Diseases”

  1. I studied pure maths (cryptography + number theory)
  2. I got a book about evolution, apes and genes when I was in my 4th year

5 of 72

“how it came about that you are working in Rare Diseases”

6 of 72

“how it came about that you are working in Rare Diseases”

  1. Refused a job in Russian secret service FSB.
  2. Got offers from some banks and IT companies
  3. I tried to get a PhD position in bioinformatics but failed
  4. Started MSc program in algorithmic bioinformatics
  5. The first semester project was from a company providing NGS-based neonatal screening (CFTR, PAH, GALT)

7 of 72

Short introduction

  1. In 1869, Friedrich Miescher isolated "nuclein," DNA with associated proteins, from cell nuclei. 

Image credit: https://www.historischer-augenblick.de/dna-schlosskueche/

8 of 72

Genetic variants

Image credit: Nesta, Alex & Tafur, Denisse & Beck, Christine. (2020). Hotspots of Human Mutation. Trends in Genetics. 37. 10.1016/j.tig.2020.10.003.

9 of 72

Structural variants mechanisms

Image credit: Currall, B.B., Chiangmai, C., Talkowski, M.E. et al. Mechanisms for Structural Variation in the Human Genome. Curr Genet Med Rep 1, 81–90 (2013). https://doi.org/10.1007/s40142-013-0012-8

10 of 72

Human genome reference: hg19 vs hg38 vs T2T vs pan-genome

Image credit: https://www.nist.gov/image/human-reference-genome-old-vs-new

11 of 72

T2T vs hg38 Genome

Image credit: T2T Paper, Science

12 of 72

Pan-genome�(MC – Minigraph-Cactus,�PGGB – GraphBuilder)

Image credit: A draft human pangenome reference

13 of 72

Sequencing technologies

Image credit: https://www.pacb.com/blog/the-evolution-of-dna-sequencing-tools/

14 of 72

[Whole] Genome vs Exome vs Targeted NGS

Image credit: https://www.novogene.com/eu-en/resources/blog/a-beginners-guide-to-dna-sequencing-blog/

15 of 72

Alignment: BWA MEM, BWA SW, BowTie2, etc

Image credit: https://rbatorsky.github.io/intro-to-ngs-bioinformatics/lessons/03_Alignment.html, DOI:10.1109/IPDPS.2019.00041

16 of 72

Repeats

Black – unmasked

Dark blue/blue/Purple – L1/L2/SINE(alu)

Image credit: https://www.repeatmasker.org/species/hg.html

17 of 72

Transposable elements in human genomes

Image credit: https://www.nature.com/articles/s41576-019-0165-8

18 of 72

Paralogous genes

Image credit: Recent Progress in Gene-Targeting Therapies for Spinal Muscular Atrophy: Promises and Challenges

19 of 72

Complex rearrangements

  1. Image credit: αβ T-cell receptor bias in disease and therapy (Review)

20 of 72

Non problematic regions…

  1. In this workshop we will deal with the “non-problematic for short reads” regions only, which means: short read, sequenced from this part of the human genome, can be confidently mapped to the reference genome
  2. It still does not mean the variants we will see will be error-free – manual inspection is necessary

21 of 72

Characteristics of [short] variants

  1. Somatic (mosaic) or germline
  2. Homozygous, hemizygous or heterozygous [compound heterozygous]
  3. True positive, false positive, false negative
  4. Coding and non coding
  5. Substitution, insertion/deletion (<50bp) – indels
  6. Maternally/paternally inherited or de novo
  7. Missense, protein-truncating (frameshift, premature stop, etc), splice variants. VEP annotation (HIGH/MODERATE/MODIFIER/LOW impact)
  8. Pathogenic (class 5), likely pathogenic (class 4), VUS, benign (also for long variants)
  9. Causal variant, partially explaining phenotype, non relevant (also long)

22 of 72

Variant classification

PMID: 25741868

23 of 72

Variant classification

Source: Navigating the nuances of clinical sequence variant interpretation in Mendelian disease

24 of 72

Scoring “on the piece of paper”

25 of 72

Variant classification

26 of 72

Evaluation of variants is nuanced collecting points

27 of 72

Variant quality evaluation

  1. Coverage – variant detected in a region covered with 10 reads has a different evidence than the one found in a region with a coverage of 100x
  2. Strand bias – e.g. for a particular position A (25+,20-), C(35-,25+) – seems fine, A(25+,0-), C(35-,25+) – not.
  3. PCR artifacts – if you have multiple copies of the same read (start/end usually match) and your variant is presented only in these reads
  4. For indels – the borders should be “clean”, nicely aligned, soft-clipped parts – representing the same sequence
  5. Presence of a lot of soft-clipped sequences around the variant is a bad sign too

28 of 72

Variants longer than 50bp?

  1. CNV
  2. STR expansion
  3. Mobile Element Insertion
  4. SV (includes CNVs, MEIs, STRs)

Img credit: Structural variant identification and characterization

29 of 72

Aneuploidy

30 of 72

Uniparental disomy

  1. Image credit: Diagnostic testing for uniparental disomy: a points to consider statement from the American College of Medical Genetics and Genomics (ACMG)

31 of 72

Methods for CNV detection

  1. Locus-specific
  2. Genome-wide

32 of 72

RT-qPCR

  1. Parental somatic mosaicism for CNV deletions – A need for more sensitive and precise detection methods in clinical diagnostics settings

33 of 72

MLPA

  1. MLPA-Based Analysis of Copy Number Variation in Plant Populations

34 of 72

MLPA

  1. Case of chromosome 22q11.2 deletion syndrome in Russian family

35 of 72

Microarrays

36 of 72

B-allele frequency: 100x coverage of position X

50 reads show C, 50 reads show A – 0.5 or 50%

51 reads show C, 49 reads show A – 0.49 or 49%

DELETION

50 reads show C, 0 reads show A – 1.0 or 100%

DUPLICATION

50 reads show C, 100 reads show A – 0.33 or 33%

TRIPLICATION

50 reads show C, 150 reads show A – 0.25 or 25%

37 of 72

Methods for CNV detection

  1. Detecting copy number variation in next generation sequencing data from diagnostic gene panels

38 of 72

Paired-end sequencing

  1. reference: https://www.ecseq.com/support/ngs/why-do-the-reads-all-have-the-same-length-when-sequencing-differently-sized-fragments

39 of 72

Methods for SV detection

  1. reference: https://www.researchgate.net/publication/27505392

40 of 72

[Whole] Genome vs Exome vs Targeted NGS

Image credit: https://www.novogene.com/eu-en/resources/blog/a-beginners-guide-to-dna-sequencing-blog/

41 of 72

Exome kits

42 of 72

Agilent v5�3.5K samples

43 of 72

IGV color scheme (needs to be activated)

  1. reference: IGV website

44 of 72

IGV color scheme

  1. reference: IGV website

45 of 72

IGV color scheme: read pair mapped to the chromosome…

  1. reference: IGV website

46 of 72

Bioinformatics Behind: Read Depth Based CNV calling

  1. Divide reference genome into intervals
  2. Calculate read coverage in these intervals
  3. Find stretches of intervals with unusually high or low coverage

47 of 72

Low/high coverage are signs of CNVs?

  1. https://pmbio.org/module-03-align/0003/04/01/PostAlignment_Visualization/

48 of 72

Low/high coverage are signs of CNVs?

  1. https://pmbio.org/module-03-align/0003/04/01/PostAlignment_Visualization/

49 of 72

GC Bias

Source: Summarizing and correcting the GC content bias in high-throughput sequencing

50 of 72

XHMM

Source: Discovery and Statistical Genotyping of Copy-Number Variation from Whole-Exome Sequencing Depth

51 of 72

Circular Binary Segmentation

52 of 72

Modern tools

Open source or commercial? WES/WGS?

Open source: CNV-kit, Conifer, ExomeDepth, ClinCNV

Commercial: QUIAGEN tool, SeqNext (JSI), DRAGEN (Illumina)

53 of 72

CNV classification

54 of 72

CNV classification

55 of 72

Franklin

  1. Source: https://franklin.genoox.com/clinical-db/variant/sv/chr1-216138600-216270555-DEL

56 of 72

Franklin

  1. Source: https://franklin.genoox.com/clinical-db/variant/sv/chr1-216138600-216270555-DEL

57 of 72

Examples

  1. The following examples are from the real WES samples, submitted within Solve-RD under broad research consent. Some of the submitted samples belong to patients, some to their healthy relatives. Causal variants will be submitted to ClinVar.
  2. These pictures DO NOT allow the re-identification of the samples. The screenshots are fully annonymized so even I don’t know any details.
  3. However, please, do not spread the materials further.

58 of 72

IGV color scheme: read pair mapped to the chromosome…

  1. Solve-RD project

59 of 72

IGV color scheme: read pair mapped to the chromosome…

  1. Solve-RD project

60 of 72

IGV color scheme: read pair mapped to the chromosome…

  1. Solve-RD project

61 of 72

62 of 72

63 of 72

CNV annotation

  1. Tools for comprehensive annotation such as AnnotSV – a lot of information, not adapted for your internal database
  2. In house solution – difficult to implement same info fields, but calling is identical to yours
  3. Web service interpretation – external resources with broader descriptions

64 of 72

AnnotSV web version

65 of 72

AnnotSV full TSV

AnnotSV_ID SV_chrom SV_start SV_end SV_length SV_type Samples_ID Annotation_mode CytoBand Gene_name Closest_left Closest_right Gene_count Tx Tx_version Tx_start Tx_end Overlapped_tx_length Overlapped_CDS_length Overlapped_CDS_percent Frameshift Exon_count Location Location2 Dist_nearest_SS Nearest_SS_type Intersect_start Intersect_end RE_gene P_gain_phen P_gain_hpo P_gain_source P_gain_coord P_loss_phen P_loss_hpo P_loss_source P_loss_coord P_ins_phen P_ins_hpo P_ins_source P_ins_coord po_P_gain_phen po_P_gain_hpo po_P_gain_source po_P_gain_coord po_P_gain_percent po_P_loss_phen po_P_loss_hpo po_P_loss_source po_P_loss_coord po_P_loss_percent P_snvindel_nb P_snvindel_phen B_gain_source B_gain_coord B_gain_AFmax B_loss_source B_loss_coord B_loss_AFmax B_ins_source B_ins_coord B_ins_AFmax B_inv_source B_inv_coord B_inv_AFmax po_B_gain_allG_source po_B_gain_allG_coord po_B_gain_someG_source po_B_gain_someG_coord po_B_loss_allG_source po_B_loss_allG_coord po_B_loss_someG_source po_B_loss_someG_coord GC_content_left GC_content_right Repeat_coord_left Repeat_type_left Repeat_coord_right Repeat_type_right Gap_left Gap_right SegDup_left SegDup_right ENCODE_blacklist_left ENCODE_blacklist_characteristics_left ENCODE_blacklist_right ENCODE_blacklist_characteristics_right ACMG HI TS DDD_HI_percent ExAC_delZ ExAC_dupZ ExAC_cnvZ ExAC_synZ ExAC_misZ GenCC_disease GenCC_moi GenCC_classification GenCC_pmid NCBI_gene_ID OMIM_ID OMIM_phenotype OMIM_inheritance OMIM_morbid OMIM_morbid_candidate LOEUF_bin GnomAD_pLI ExAC_pLI AnnotSV_ranking_score AnnotSV_ranking_criteria ACMG_class

66 of 72

GSvar

67 of 72

In House

68 of 72

DDD

69 of 72

DDD

70 of 72

DDD

71 of 72

DDD

72 of 72

Vielen Dank für

Ihre Aufmerksamkeit