Development of a high-performance processing pipeline for the discovery of variants interaction and its association with complex disorders
Why extreme data?
*MN4 architecture: 165,888 CPU - 1.880 GB/core and 208 GPU -16GB/each.
1. Multi Dimensionality Reduction (MDR)
MDR stages.
MDR uses statistical methods to discover pairs of variants
which, synergically, contribute to the development of Type 2 Diabetes (T2D)
2. Genome-Wide discovery (GWD)
GWD stages.
GWD uses machine learning methods to find groups of variants that are simultaneously associated with T2D.
Components integrated
MPI The Message Passing Interface (MPI) is an Application Program Interface that defines a model of parallel computing where each parallel process has its own local memory, and data must be explicitly shared by passing messages between processes.
Lithops is a distributed computing framework for data analysis at massive scale that fits perfectly into highly parallelizable programs without the need for inter-process communication, but also supports parallel applications that need to share state between processes.
Multi Dimensionality Reduction (MDR) and Genome-Wide discovery (GWD)
OPEN MPI
Datasets
70,127 individuals
12,931 diabetic
57,196 non-diabetic
15,131,345 variants
70K dataset
(Bonàs-Guarch et al., 2018)
Confidentiality constraints.
UKBioBank dataset
(Sudlow et al., 2015)
Confidentiality constraints.
422,000 individuals
31,344 diabetic
233,285 non-diabetic
15,586,493 variants
Synthetic dataset
Non confidentiality constraints.
Development purposes
n individuals
m diabetic
p non-diabetic
x variants
Multi Dimensionality Reduction (MDR) and Genome-Wide discovery (GWD)
KPIs
Multi Dimensionality Reduction (MDR)
KPI-1: New MPI version is 4x times faster than first Spark version processing of all variants
�
Projection for 15 million variants, involving 6x10¹³ combinations, utilizing RES* infrastructure.
*RES: Red Española de Supercomputación
COMBS/SEC/CORE
MN4-Python: 507
MN4-MPI: 1311
MN5-MPI: 2000
KPI-1: Significant performance improvements (data throughput, data transfer reduction) in
Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data
volumes (genomics, metabolomics).
KPIs
Genome-Wide discovery (GWD)
KPI-1: Significant performance improvements (data throughput, data transfer reduction) in
Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data
volumes (genomics, metabolomics).
KPI-3: Demonstrated resource auto-scaling for batch and stream data processing, validated
thanks to data-driven orchestration of massive workflows.
KPI1 & KPI-3: Lithops-HPC architecture aims to offer high performance and scalability while simplifying application coding and deployment
Lithops client
N functions
Machine 1
Function 2
Function 3
Function 1
Function N
Singularity container
Function 2
Function 3
Function 1
Function N
Function 2
Function 3
Function 1
Function N
GPFS Storage
Lithops' benchmarks on MareNostrum5 outperform those of other commercial cloud platforms in terms of GFLOPS (left fig) and read-write bandwidth (right fig).
Machine 2
Machine n
Singularity container
Singularity container
Demo video
Multi Dimensionality Reduction (MDR)
MN5 SLURM scheduling
MN5 terminal
Demo video
Genome-Wide discovery (GWD)
MN5 SLURM scheduling
Lithops backend
Client interface
Publications
T5.1 AI-Based Optimizations
MDR use-case
Initial analysis using Kuberntes over a bare-metal cluster and MDR use-case.
Predict resource usage for the next time window (parameter)