1 of 54

Analysis of SARS-CoV-2 Sequencing Data

Alex Zelikovsky

Department of Computer Science

Georgia State University

​

CGSI 2022

​

July 15, 2022

2 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
  • Network Representation of SARS-CoV-2 Sequences: eMST

​

​

3 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
  • Network Representation of SARS-CoV-2 Sequences: eMST

​

​

4 of 54

https://www.nature.com/articles/s41592-022-01444-z

5 of 54

6 of 54

Genomics-based methods enabled early warnings of COVID-19 pandemic

Illumina.com

7 of 54

Mobilization of sequencing resources

8 of 54

Sequencing platforms

9 of 54

Disparities in SARS-CoV-2 sequencing

10 of 54

Disparities in SARS-CoV-2 sequencing

11 of 54

Sequencing in Africa

$843/capita

$800/capita

$4,100/capita

$3,700/capita

12 of 54

Main contribution of genomics and bioinformatics

  • Enabled early warnings of COVID-19 pandemic
  • Shaped the effective COVID-19 response
  • Enabled tracking COVID-19 geographical spread in real time
  • Tracked SARS-CoV-2 genomic evolution
  • Unlocked wastewater-based monitoring

(Grubaugh et al. 2019)

13 of 54

During the COVID-19 pandemic, genomics and bioinformatics have emerged as essential public health tools.

14 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
  • Network Representation of SARS-CoV-2 Sequences: eMST

​

​

15 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
    • RNA Haplotypes
    • CliqueSNV
    • CliqueSNV-based Clustering
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
  • Network Representation of SARS-CoV-2 Sequences: eMST

​

16 of 54

17 of 54

RNA-virus replication => Haplotypes

​

​

High mutation rate (~10-4)

​

​

​

​

​

​

Haplotypes differ in

- Virulence

- Escape immune response

- Resistance to antiviral therapies

Haplotypes detection largely relevant for HIV and HCV

18 of 54

Intra-host viral population analysis

​

​

Given:

NGS reads from viral sample.

Find:

  1. The number of different haplotypes in the sample.
  2. The frequencies of these haplotypes.
  3. The genome of each haplotype.

Obstacles:

  1. We do not know which of the haplotypes the reads were generated from.
  2. The length of a read is much shorter than the haplotypes’ length.
  3. The reads are error-prone, thus a read string may not exactly match that part of the haplotype it was sequenced from.
  4. The starting positions of the reads with respect to the haplotypes are unknown.

19 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
    • RNA Haplotypes
    • CliqueSNV
    • CliqueSNV-based Clustering
  • Network Representation of SARS-CoV-2 Sequences: eMST
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE

​

20 of 54

21 of 54

22 of 54

Linked SNV pairs Forbidden SNV pairs�

O22 = # of observed minor allele co-occurrences in reads

T22 = frequency of true minor allele co-occurrences in reads

t = minimum detectable haplotype frequency

empirical default =0.1%

23 of 54

24 of 54

CliqueSNV performance

Competitors:

  • PredictHaplo (Prabhakaran et al., 2013)
  • 2SNV (Artyomenko et al. 2017)
  • aBayesQR (Ahn and Vikalo, 2018)

Datasets:

25 of 54

CliqueSNV performance

26 of 54

Earth Mover's Distance

27 of 54

Impact

CliqueSNV was used to analyze evolution events in

​

  • COVID-19 patients:
    • Ringlander J, et al. Impact of ADAR-induced editing of minor viral RNA populations on replication and transmission of SARS-CoV-2. PNAS 2022
    • Gaiarsa S, et al. Comparative analysis of SARS-CoV-2 quasispecies in the upper and lower respiratory tract shows an ongoing evolution in the spike cleavage site. Virus Res. 2022 
  • HIV patients:
    • Kemp, et al. "HIV-1 Evolutionary Dynamics under Non-suppressive Antiretroviral Therapy." Mbio (2022)
  • animals:
    • Fish, et al. "Foot-and-Mouth Disease Virus Interserotypic Recombination in Superinfected Carrier Cattle. Pathogens 2022, 11, 644." (2022).

28 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
    • RNA Haplotypes
    • CliqueSNV
    • CliqueSNV-based Clustering
  • Network Representation of SARS-CoV-2 Sequences: eMST
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE

​

29 of 54

Given GISAID data

  • Find subtypes/variants by clustering
  • Assess clustering quality
  • Estimate fitness of each cluster over time period

​

30 of 54

Clique SNV-based clustering

Idea:

    • Align sequences all/part of GISAID sequences
    • Run Clique SNV treating them
      • as if they are coming from single intra-host viral population
      • as if they are whole-genome PacBio reads of good quality
    • The output haplotypes can be treated as SARS-C0V-2 subtypes/variants
    • Cluster GISAID sequences by attributing to the closest haplotype

31 of 54

Assessing clustering of categorical data

​

Sequences = categorical data

Standard approaches:

    • Embed in a metric space
      • Estimate distance between data points
        • Hamming distance
        • TN93 (not even a real distance)
      • Cluster points in a metric space
    • Embed in multidimensional space and use any powerful clustering tools

Drawbacks

– distance introduces error

– should be position-dependent …

Idea: Approach from

Li, T., Ma, S., and Ogihara, M. 2004. Entropy-based criterion in categorical clustering.

​

​

32 of 54

Clustering Entropy

​

​

​

​

​

33 of 54

Gap-filling algorithms

    • Consensus-based
      • Obtain consensus for the entire dataset:
        • Align GISAID sequences
        • Find consensus
      • For each sequence, fill gaps using consensus

​

    • Reference-based
      • For each sequence, fill gaps using reference

​

    • Clustering-based
      • For each sequence, fill gaps using cluster consensus

​

​

​

34 of 54

Entropy-based gap filling and clustering assessment

​

​

k-modes setting(initialization, distance)

Without gap filling

With gap filling

Expected entropy

Total

entropy

Expected entropy

Total

entropy

without clustering

9536.89

9536.89

8417.89

8417.89

random centers, Hamming

123.00

3170.60

109.21

2474.30

random centers, TN-93

127.32

4401.18

111.05

3470.03

pairwise distant, Hamming

422.65

4651.23

294.98

3629.47

pairwise distant, TN-93

273.34

3500.14

256.44

3007.07

Clique SNV, Hamming

110.58

2585.29

90.42

2308.95

Clique SNV, TN-93

121.87

2379.46

100.85

2117.40

35 of 54

Haplotype distribution (GISAID dataset, cumulative, relative count)�

36 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
    • Just make it no backward mutations
    • Runtime
  • Network Representation of SARS-CoV-2 Sequences: eMST

​

​

37 of 54

  • Parsimony character-based method
  • Allow recurrent mutations
  • Forbid backward mutations
  • Assign seq’s to all internal nodes
  • Scalable: 200K GISAID seqs in <2h

38 of 54

Tracking of SARS-CoV-2 evolution in real time

(Fauver et al. 2020)

39 of 54

Parsimony resurged in viral phylogenetics

40 of 54

Problem Formulation

  • Given:
    • A set of aligned sequences, (possibly containing missing positions)
    • A reference sequence with no missing positions
  • Find: a character-based phylogenetic tree such that:
    • Rooted at the reference sequence
    • Has the minimum total edge length
    • No backward mutations are allowed
      • This is specific to our model in this presentation
  • Perfect phylogeny (no backward or recurrent mutations)⬄
    • Hamming distance b/w any haplotypes = distance along the tree
  • Backward mutations not allowed, but recurrent mutations allowed ⬄
    • Hamming distance b/w ancestor and descendent = distance along the tree

41 of 54

Algorithm description

Input: Aligned sequences (with possible missing positions), reference sequence

Output: Character-Based Phylogenetic Tree on aligned sequences, rooted at reference sequence.

Backward mutations are not allowed, recurrent mutations are allowed.

​

1. Initialize a tree with reference as root, and a queue of all sequences

2. Set priority for each enqueued sequence = Hamming distance to root

3. Set parent for each enqueued sequence = root

​

4. While the queue is not empty:

​

5. Remove minimum priority sequence x

​

6. for each node p in reversed(tree):

check if p is the parent of x

stop when a parent is found

​

7. if Hamming distance to x’s parent is 0:

collapse x to its parent

​

else add x to the tree:

connect x to its parent

fill missing positions in x from parent haplotype

42 of 54

SPHERE Runtime

Runtime is subquadratic

​

Experimental Estimate O(n^ 1.6)

43 of 54

Comparison with NextStrain on Coast-to-Coast

44 of 54

Stability: SPHERE vs NextStrain

  • Directional normalized Robinson-Foulds distance
    • Bipartitions present in row tree not present in column tree
    • Normalized by number of bipartitions in the row tree

Consider:

45 of 54

Plan

  • Genomics response to the COVID-19 pandemics
  • Applying CliqueSNV to SARS-CoV-2 sequences
  • Scalable SARS-CoV-2 phylogeny reconstruction: SPHERE
  • Network Representation of SARS-CoV-2 Sequences: eMST
    • Combination of multiplicative and absolute constraints
    • Assortativity

​

​

46 of 54

A Novel Network Representation of SARS-CoV-2 Sequencing Data

Sergey Knyazev, Daniel Novikov, Mark Grinshpon, Harman Singh, Ram Ayyala, Varuni Sarwal, Roya Hosseini, Pelin Icer Baykal, Pavel Skums, Ellsworth Campbell, Serghei Mangul & Alex Zelikovsky

 

47 of 54

Network Representation of SARS-CoV-2 Sequences: eMST

 

48 of 54

Validation

​

Dataset: Early Transmission Links

  • 293 global SARS-CoV-2 sequences collected before March 9th, 2020.
  • Each sequence has a known country of origin.
  • Match the 25 known country-to-country transmission links

​

Attribute Assortativity is a preference for a network's nodes to attach to others that are similar w.r.t the attribute

​

​

49 of 54

 

50 of 54

 

51 of 54

 

52 of 54

Summary (repeating …)�During the COVID-19 pandemic, genomics and bioinformatics have emerged as essential public health tools.

53 of 54

HIV Subdivision

  • Bill Switzer
  • Ellsworth Campbell

​

Hepatitis Subdivision

  • Yury Khudyakov
  • David S. Campo

​

​

Computer Science Department

  • Pavel Skums
  • Murray Patterson
  • Sarwan Ali
  • Sergey Knyazev ( UCLA/USC)
  • Vyacheslav Tsyvina (Facebook)
  • Andrew Melnyk(Firework Studios)
  • Pelin I. Baykal (ETH Zurich)
  • Daniel Novikov
  • Roya Hosseini
  • Fatemeh Mohebbi
  • Bikram Sahoo
  • Mark Grinshpon

NSF Grant DBI-1564899

NSF Grant CCF-1619110

NIH Grant 1R01EB025022-01

GSU Molecular Basis of Disease

Fellowship

  • Serghei Mangul
  • Karishma Chhugani
  • Ram Ayyala
  • Harman Singh
  • Varuni Sarwal

54 of 54

NSF (travel fellowships)

ISCB (member discount)

Georgia State University

University of Haifa

University of South California

18th Intl Symposium on Bioinformatics Research & Applications�

Full paper – proceedings 07/25, 2022

Full paper - journal invitation 07/25, 2022

Notifications 08/25, 2022�Short papers 08/27, 2022

Highlights 09/07, 2022

Breakthrough 09/07, 2022

HAIFA, ISRAEL

Tracks:

Sponsors:

Special�Issues:

IEEE/ACM TCBB

Journal of Computational Biology

BMC Genomics

BMC Bioinformatics