1 of 13

Development of a high-performance processing pipeline for the discovery of variants interaction and its association with complex disorders

2 of 13

Why extreme data?

  • Variants can be associated with complex diseases.
  • Using the MareNostrum 4 (MN4) architecture* to study all possible combinations of pairs (6x10¹³ tests) takes approximately 9 days.
    • Analyzing all possible combinations across the 23 chromosomes would require months of computation, even with the entire supercomputer.
  • Two use cases are being conducted to study variant interactions in Type 2 Diabetes (T2D).

​

​

*MN4 architecture: 165,888 CPU - 1.880 GB/core and 208 GPU -16GB/each.

3 of 13

1. Multi Dimensionality Reduction (MDR)

MDR stages.

  1. k-fold Cross-Validation: Divide the dataset into k-1 parts for training and 1 for test, repeating k times.
  2. Genotype Cross-Tabulation: Create 9-cell tables for each variant pair from training data and count the number of cases and controls.
  3. 1-Dimension Transformation: Classify cells as ‘high risk’ or ‘low risk’ using a case-control ratio threshold.
  4. Individual Classification: Classify individuals based on table cells and compute miss-classification error using the test part.
  5. Model Selection: Choose the model with the best error rate and evaluate its predictive power, selecting consistent variant pairs.

MDR uses statistical methods to discover pairs of variants

which, synergically, contribute to the development of Type 2 Diabetes (T2D)

4 of 13

2. Genome-Wide discovery (GWD)

GWD stages.

  1. Read Stage: The dataset containing the genetic variants is read and prepared for analysis.
  2. Train Stage: A machine learning model (i.e., XGBoost) is trained using the prepared dataset. This involves:
    1. Hyperparameter Tuning.
    2. Data Partitioning.
    3. Model Training.
  3. Test Stage: The trained model is evaluated using the test set to measure its performance.

GWD uses machine learning methods to find groups of variants that are simultaneously associated with T2D.

5 of 13

Components integrated

MPI The Message Passing Interface (MPI) is an Application Program Interface that defines a model of parallel computing where each parallel process has its own local memory, and data must be explicitly shared by passing messages between processes.

​

Lithops is a distributed computing framework for data analysis at massive scale that fits perfectly into highly parallelizable programs without the need for inter-process communication, but also supports parallel applications that need to share state between processes.

Multi Dimensionality Reduction (MDR) and Genome-Wide discovery (GWD)

OPEN MPI

6 of 13

Datasets

70,127 individuals

12,931 diabetic

57,196 non-diabetic

15,131,345 variants

70K dataset

(Bonàs-Guarch et al., 2018)

Confidentiality constraints.

UKBioBank dataset

(Sudlow et al., 2015)

Confidentiality constraints.

422,000 individuals

31,344 diabetic

233,285 non-diabetic

15,586,493 variants

Synthetic dataset

Non confidentiality constraints.

Development purposes

n individuals

m diabetic

p non-diabetic

x variants

Multi Dimensionality Reduction (MDR) and Genome-Wide discovery (GWD)

7 of 13

KPIs

Multi Dimensionality Reduction (MDR)

​

​

KPI-1: New MPI version is 4x times faster than first Spark version processing of all variants

�

Projection for 15 million variants, involving 6x10¹³ combinations, utilizing RES* infrastructure.

*RES: Red Española de Supercomputación

COMBS/SEC/CORE

MN4-Python: 507

MN4-MPI: 1311

MN5-MPI: 2000

KPI-1: Significant performance improvements (data throughput, data transfer reduction) in

Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data

volumes (genomics, metabolomics).

8 of 13

KPIs

Genome-Wide discovery (GWD)

​

KPI-1: Significant performance improvements (data throughput, data transfer reduction) in

Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data

volumes (genomics, metabolomics).

​

KPI-3: Demonstrated resource auto-scaling for batch and stream data processing, validated

thanks to data-driven orchestration of massive workflows.

KPI1 & KPI-3: Lithops-HPC architecture aims to offer high performance and scalability while simplifying application coding and deployment

Lithops client

N functions

Machine 1

Function 2

Function 3

Function 1

Function N

Singularity container

Function 2

Function 3

Function 1

Function N

Function 2

Function 3

Function 1

Function N

GPFS Storage

Lithops' benchmarks on MareNostrum5 outperform those of other commercial cloud platforms in terms of GFLOPS (left fig) and read-write bandwidth (right fig).

Machine 2

Machine n

Singularity container

Singularity container

9 of 13

Demo video

Multi Dimensionality Reduction (MDR)

​

​

  1. A configuration file is employed to generate a synthetic dataset comprising 5,600 variants, distributed across 28 files, each containing 200 samples.
  2. MPI is utilized to compute 15.6 million variant combinations (5,600 × 5,600 / 2), distributing the task across 384 CPUs.
  3. The execution results in 384 files, each containing 40,000 combinations, and completes in approximately 20 seconds.

MN5 SLURM scheduling

MN5 terminal

10 of 13

Demo video

Genome-Wide discovery (GWD)

​

​

  1. The Lithops backend initiates with 100 workers (1 cluster node utilizing 100 CPUs).
  2. The RabbitMQ server and Lithops workers are set up to manage client deployments.
  3. The client deploys the FLOPS benchmark, which is then processed by the backend.

MN5 SLURM scheduling

Lithops backend

Client interface

11 of 13

Publications

  • Gómez-Sánchez, G.; Alonso, L.; Pérez, M.Á.; Morán, I.; Torrents, D.; Berral, J.L. Exhaustive Variant Interaction Analysis Using Multifactor Dimensionality Reduction. Appl. Sci. 2024, 14, 5136. https://doi.org/10.3390/app14125136

​

  • Gonzalo Gómez-Sánchez, Aaron Call, Xavier Teruel, Lorena Alonso, Ignasi Moran, Miguel Ángel Perez, David Torrents, Josep Ll. Berral. Challenges and Opportunities for RISC-V Architectures towards Genomics-based Workloads. First International workshop on RISC-V for HPC, of ISC High Performance conference, 2023.

​

  • Enhancing HPC with Serverless Computing: Lithops on MareNostrum5. Andrés Benavides, Daniel Coll-Tejeda, Aaron Call, Pedro García, Ramon Nou. In Cloud-Edge Continuum (CEC) Workshop 2024, Co-located with The 32nd IEEE International Conference on Network Protocols. October 2024, Charleroi, Belgium.

​

  • Received RES grant ID BCV-2024-2-0004, “Extreme-data processing platforms to analyze the interaction of genomic variants and their association to common diseases”. Aaron Call, Lorena Alonso, Andrés Benavides, Ramon Nou.
    • 9 million CPU hours of prioritary usage in MareNostrum 5 to run NEARDATA experiments. July 1st, 2024.

12 of 13

T5.1 AI-Based Optimizations

MDR use-case

​

Initial analysis using Kuberntes over a bare-metal cluster and MDR use-case.

​

  • The more steps ahead are better to compare.
  • NN model used here: we also tested with linear models.
  • To be integrated into Lithops.

Predict resource usage for the next time window (parameter)

13 of 13