1 of 16

Genomics Use-Case

Leader: UK Health Security Agency (UKHS)

Collaboration: Universitat Rovira i Virgili (URV)

EXTREME NEAR-DATA PROCESSING PLATFORM

2 of 16

Why is the use case extreme data?

  • Genome sequence analysis is a compute- and data-intensive task.
    • Exponential growth in data size and complexity.
    • Biomedical institutions with HPC struggle to keep up.

​

  • Genomics is a key use-case in extreme data analytics, due to its increasing volume, variety and complexity.

​

  • Use case for genomic sequence analysis based on variant calling to identify variants from reference genome sequences.

​

3 of 16

Assessed Datasets

  • Small size:
    • Trypanosome Reference genome: TriTrypDB-67_TbruceiTREU9271 (35MB).
    • Sequence Reads: SRR60521332 (668MB).

​

  • Medium size:
    • Human Reference genome: hg193 (905MB).
    • Sequence Reads: SRR150683234(1.2GB) and ERR98564895 (12.1 GB).

​

  • Large size:
    • Bos taurus Reference genome: bos_taurus6 (781MB).
    • Sequence Reads: SRR9344157 (16.5GB).

Two genomics data formats:

    • FASTA
    • FASTQ

4 of 16

Description of the use-case

  • Objective: Adapt an existing single-node HPC variant calling genomics application to serverless in order to scale in parallelism, process larger datasets and decrease runtime.

​

  • Variant Calling: detect differences (mutations, variants) in a sampled genome compared to a reference genome.

​

5 of 16

Serverless Variant Calling workflow

6 of 16

Data Connectors

  • Data Partitioner: Dataplug

7 of 16

Data Connectors

  • Data Loader: Lithops Storage API

8 of 16

Data Connectors

  • Data Shuffling: AWS S3 Select

9 of 16

Data Connectors

  • New Near-Data Shuffling: Glider

10 of 16

Data Connectors

  • Data Merger: AWS Multipart Upload

11 of 16

What are the benchmarks/KPI?

  • KPI-1 - Significant performance improvements (data throughput, data transfer reduction) in Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data volumes (genomics).

​

  • Dataplug offers up to a 4x improvement in execution time and reduces data transfers by 200%.

12 of 16

What are the benchmarks/KPI?

  • KPI-3 - Demonstrated resource auto-scaling for batch and stream data processing validated thanks to data-driven orchestration of massive workflows.

​

  • Lithops: The serverless version is x37.46 times faster than the HPC version.

13 of 16

What are the benchmarks/KPI?

  • KPI-3 - Demonstrated resource auto-scaling for batch and stream data processing validated thanks to data-driven orchestration of massive workflows.

​

  • Lithops: Scalability experiment with the medium-sized dataset.

14 of 16

What are the benchmarks/KPI?

  • KPI-3 - Demonstrated resource auto-scaling for batch and stream data processing validated thanks to data-driven orchestration of massive workflows.

​

  • Glider: Scalability and performance experiment with the medium-sized dataset (Reduce phase).
    • Glider reduces execution time by 36% with the full data.
    • Glider demonstrates a reduction of the data transfer ensuring the workflow orchestration from the distribution of the actions.

15 of 16

Github repository

  • https://github.com/neardata-eu/serverless-genomics-variant-calling

16 of 16

Thank you