1 of 19

Overview and results

Leader: Universitat Rovira i Virgili (URV)

EXTREME NEAR-DATA PROCESSING PLATFORM

2 of 19

List of participants

3 of 19

HORIZON-CL4-2022-DATA-01-05: �Extreme data mining, aggregation and analytics technologies and solutions

Provide better technologies, tools and solutions for data mining (searching and processing) of extreme data.

​

Extreme data is defined as data that exhibits one or more of the following characteristics, to an extent that makes current technologies fail: increasing volume, speed, variety; complexity/diversity/multilinguality of data; the dispersed data sources; sparse/missing/insufficient data/extreme variations in values).

​

The technologies and solutions are expected to discover and distil meaningful, reliable and useful data from heterogeneous and dispersed/scarce sources and deliver it to the requesting application/user with minimal delay and in the appropriate format.

4 of 19

Extreme Near-Data Processing Platform

Why Locality ?

Volume

Privacy

Low latency

Hardware Acceleration

​

5 of 19

Objectives

  • O-1 Provide high-performance near-data processing for Extreme Data Types: The first objective is to create a novel intermediary data service (Data Plane / Xtreme DataHub) between Object Storage and Analytic platforms.
  • O-2 Support real-time video streams but also event streams that must be ingested and processed very fast to Object Storage: The second objective is to seamlessly combine streaming and batch data processing for analytics.
  • O-3 Provide secure data orchestration, transfer, processing and access: The third objective is to create a Data Broker service enabling trustworthy data sharing and confidential orchestration of data pipelines across the Compute Continuum.

The main goal is to design an Extreme near-data processing platform to enable consumption, mining and processing of distributed and federated data without needing to master the logistics of data access across heterogeneous data locations and pools.

6 of 19

KPIs

  • KPI-1 - Significant performance improvements (data throughput, data transfer reduction) in Extract-Transform-Load (ETL) phases validated with near-data connectors over extreme data volumes (genomics, metabolomics).

​

  • KPI-2 - Significant data speed improvements (throughput, latency) in real-time video analytics validated using stream data connectors.

​

  • KPI-3 - Demonstrated resource auto-scaling for batch and stream data processing validated thanks to data-driven orchestration of massive workflows.

​

  • KPI-4 - High levels of data security and confidential computing validated using TEEs and federated learning in adversarial security experiments.

​

  • KPI-5 - Demonstrated simplicity and productivity of the software platform validated with real user communities in International Health Data Spaces.

7 of 19

Software components

    • Lithops: Lithops is a Python multi-cloud distributed computing framework. It allows you to run unmodified local python code at massive scale in the main serverless computing platforms. Lithops delivers the user’s code into the cloud without requiring knowledge of how it is deployed and run. Moreover, its multicloud-agnostic architecture ensures portability across cloud providers.
    • Metaspace: The METASPACE platform hosts an engine for metabolite annotation of imaging mass spectrometry data as well as a spatial metabolite knowledgebase of the metabolites from thousands of public datasets provided by the community. METASPÂCE is implemented on top of Lithops.
    • SCONE: It is a platform to build and run secure applications with the help of Intel SGX (Software Guard eXtensions). In a nutshell, the SCONE objective is to run applications such that data is always encrypted, i.e., all data at rest, all data on the wire as well as all data in main memory is encrypted. 
    • Pravega: Pravega is an open source distributed storage service implementing Streams. It offers Stream as the main primitive for the foundation of reliable storage systems: a high-performance, durable, elastic, and unlimited append-only byte stream with strict ordering and consistency.

​

8 of 19

Workpackages

9 of 19

WP2 – Global Architecture

Objectives:

​

• Design the overall system architecture of the NEARDATA software (T2.1).

​

• Provide and implement a set of interfaces and APIs to integrate the different components and software prototypes of the platform (T2.2).

​

• Describe stress-test scenarios and benchmarking framework (T2.3) and validate them in International Data Spaces (T2.4).

​

10 of 19

WP3 – Data Plane: Extreme Data Connectors

Objectives:

​

• Development of Serverless Data Connector platform for Data connectors (T3.1).

​

• Deploy Stream data connectors as stream operators (T3.2, T3.3).

​

• Validate High Performance data connectors using Hardware acceleration (T3.3).

​

• Integration with popular analytic platforms (T3.5).

Data Connectors

Results:

​

• Lithops: Resource Auto-Scaling.

​

• Pravega: Throughput and data speed improvements (x1,06 faster than Kafka and x1,4 faster than Pulsar).

​

• SCONE: Confidential Computing.

​

• Data Connectors:

      • Serverless Data Connector (Dataplug): 200% data transfer reduction and data throughput improvements (x2.9 and x3,7 faster in FASTQGZIP indexing and fetching).
      • HPC Data Connector: x5 data speed improvement compared to a cloud-based version.

11 of 19

WP4 – Control Plane: Confidential Data Orchestration

Objectives:

​

• Develop a Data broker component providing data governance and data orchestration (T4.1).

​

• Implement confidential data compute and exchange mechanisms leveraging TEEs (T4.2, T4.3).

​

• Develop confidential data orchestration mechanisms including federated learning (T4.4, T4.5).

Results:

​

• Integration of Confidential Compute Layer of Data Broker with Lithops and use-cases.

AI Component

12 of 19

WP5 – Extreme Health Use Cases

Objectives:

​

• Optimize Use Case workloads using machine learning techniques that leverage WP3 and WP4 technical achievements (T5.1).

​

• Validate the platform in Genomics, Metabolomics, and Surgery use cases with complex pipelines involving data connectors (T5.2, T5.3, T5.4, T5.5, T5.6).

​

• Create and validate open libraries of data connectors in the different use cases (T5.2, T5.3, T5.4, T5.5, T5.6).

​

  • Metabolomics

​

​

​

  • Genomics

​

​

​

  • Surgomics

​

​

13 of 19

WP5 – Extreme Health Use Cases

Metabolomics Use-case.

  • Metabolomics Use Case (EMBL):
    • Dataplug offers partitioning strategies for metabolomics data formats (ImzML).
    • Resource auto-scaling with Lithops (Datasets from under 1GB to 20GB).
    • Confidential computing with SCONE.
    • The ML-based metabolite identification is already available to users in the production version of METASPACE and is already used by METASPACE users.

​

14 of 19

WP5 – Extreme Health Use Cases

Genomics Use-cases.

  • Variant Calling Pipeline (UKHS):
    • Dataplug reduces data transfers by 200%.
    • Lithops version is x37.46 faster than HPC version.

​

  • Genomics Epistasis Use Case (BSC):
    • MPI version is x5 faster than Apache Spark version.
    • HPC Data Connector improves performance by x2.1.
    • Resource Auto-scaling with Lithops.

​

  • Transcriptomics Atlas Use Case (SANO):
    • Resource optimization reduces compute cost by 50%.
    • Confidential computing with SCONE.

​

15 of 19

WP5 – Extreme Health Use Cases

Surgery Use-case.

  • Surgery Use Case (NCT):
    • Pravega reduces end-to-end IO latency by 45%.
    • Confidential computing with SCONE.
    • Combined deployment and data management time for video analytics is reduced by 50%.

​

​

16 of 19

WP6 – Promoting impact

Objectives:

​

• Inform stakeholders (such as industry, scientific communities, EU officers, commission and administration; general public and media) about the progress of the project. A special attention will be paid to reach SMEs in involvement and utilization of the results.

​

• Encourage the production of articles, reports and demonstrations of the project results (T6.1).

​

• Perform relevant collaboration and exploitation activities (T6.2).

​

• Monitor standard setting bodies activities and contribute our achievements (T6.3).

17 of 19

Software Outcomes

18 of 19

Agenda

19 of 19

Thank you