1 of 11

WP5. Metabolomics Use-Case

2 of 11

Experiment 1

Machine Learning-based metabolite identification in METASPACE

Largest ML study in the spatial metabolomics field

Publication accepted

Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press

3 of 11

Datasets

Extreme data from METASPACE

Used datasets from METASPACE for ML

  • Training: 780 datasets
  • Test: 930 datasets
  • From 159 researchers from 47 labs

​

All datasets are available under the CC-BY 4.0 license

Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press

4 of 11

KPIs

KPI-1: Significant performance improvements (data throughput, data transfer reduction) in

Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data

volumes (genomics, metabolomics).

Scientific performance

Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press

Finding molecules �more accurately

Increased from 0.3 to 0.5

Finding more molecules

10-100 new molecules per dataset

​

Number of identified molecules

MAP (Mean Average Precision)

Comparing new ML version vs. old rule-based version

5 of 11

KPIs

KPI-1: Significant performance improvements (data throughput, data transfer reduction) in

Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data

volumes (genomics, metabolomics).

Cost

Ratio, relative to the old version

The same cost

Despite it identifies 10-100 more molecules ⇒ cheaper / molecule

Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press

6 of 11

KPIs

KPI-1: Significant performance improvements (data throughput, data transfer reduction) in

Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data

volumes (genomics, metabolomics).

Ratio, relative to the old version

Just 1.5-2 times slower, per dataset

ML identifies more molecules ⇒ approximately the same per molecule

Runtime

7 of 11

KPIs

KPI-3: Demonstrated resource auto-scaling for batch and stream data processing, validated

thanks to data-driven orchestration of massive workflows.

METASPACE-ML is implemented using Lithops http://github.com/metaspace2020/metaspace

​

Deployed on production of METASPACE http://metaspace2020.eu

​

​

​

​

​

Already used by 197 users for 2463 datasets

8 of 11

Experiment 2

Hybrid/Federated METASPACE

Building an international Data Space in metabolomics

Similar cost but speedup increases with data size. For Xenograft x2,2 faster and X089 x3,65 faster.

A hybrid version of the metabolomic pipeline has been implemented to improve its efficiency with the integration of Lithops, which uses VMs and Cloud Functions.

9 of 11

Experiment 2

Hybrid/Federated METASPACE

Federated METASPACE

Next steps: Integrate the secure version of Lithops K8s backend + SCONE with the metabolomics pipeline.

Adaptation of the hybrid version for the Lithops K8s backend.

Test performed with Brain Dataset.

10 of 11

Backup slides

11 of 11

Why extreme data?

METASPACE

50K+ datasets, 50+ TB of data

10-100 new datasets/day

Already using Lithops

A perfect use-case for extreme data technologies

Need for advanced algorithms

to reprocess 1000s & TBs of new & historic data,

in particular Machine Learning,

where it's stored (AWS),

​

yet compute-intensive!