WP5. Metabolomics Use-Case
Experiment 1
Machine Learning-based metabolite identification in METASPACE
Largest ML study in the spatial metabolomics field
Publication accepted
Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press
Datasets
Extreme data from METASPACE
Used datasets from METASPACE for ML
All datasets are available under the CC-BY 4.0 license
Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press
KPIs
KPI-1: Significant performance improvements (data throughput, data transfer reduction) in
Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data
volumes (genomics, metabolomics).
Scientific performance
Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press
Finding molecules �more accurately
Increased from 0.3 to 0.5
Finding more molecules
10-100 new molecules per dataset
Number of identified molecules
MAP (Mean Average Precision)
Comparing new ML version vs. old rule-based version
KPIs
KPI-1: Significant performance improvements (data throughput, data transfer reduction) in
Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data
volumes (genomics, metabolomics).
Cost
Ratio, relative to the old version
The same cost
Despite it identifies 10-100 more molecules ⇒ cheaper / molecule
Wadie, Stuart, et al. (2024) METASPACE-ML: Context-specific metabolite annotation for imaging mass spectrometry using machine learning, Nature Communications, in press
KPIs
KPI-1: Significant performance improvements (data throughput, data transfer reduction) in
Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data
volumes (genomics, metabolomics).
Ratio, relative to the old version
Just 1.5-2 times slower, per dataset
ML identifies more molecules ⇒ approximately the same per molecule
Runtime
KPIs
KPI-3: Demonstrated resource auto-scaling for batch and stream data processing, validated
thanks to data-driven orchestration of massive workflows.
METASPACE-ML is implemented using Lithops http://github.com/metaspace2020/metaspace
Deployed on production of METASPACE http://metaspace2020.eu
Already used by 197 users for 2463 datasets
Experiment 2
Hybrid/Federated METASPACE
Building an international Data Space in metabolomics
Similar cost but speedup increases with data size. For Xenograft x2,2 faster and X089 x3,65 faster.
A hybrid version of the metabolomic pipeline has been implemented to improve its efficiency with the integration of Lithops, which uses VMs and Cloud Functions.
Experiment 2
Hybrid/Federated METASPACE
Federated METASPACE
Next steps: Integrate the secure version of Lithops K8s backend + SCONE with the metabolomics pipeline.
Adaptation of the hybrid version for the Lithops K8s backend.
Test performed with Brain Dataset.
Backup slides
Why extreme data?
METASPACE
50K+ datasets, 50+ TB of data
10-100 new datasets/day
Already using Lithops
A perfect use-case for extreme data technologies
Need for advanced algorithms
to reprocess 1000s & TBs of new & historic data,
in particular Machine Learning,
where it's stored (AWS),
yet compute-intensive!