Overview and results
Leader: Universitat Rovira i Virgili (URV)
EXTREME NEAR-DATA PROCESSING PLATFORM
List of participants
HORIZON-CL4-2022-DATA-01-05: �Extreme data mining, aggregation and analytics technologies and solutions
Provide better technologies, tools and solutions for data mining (searching and processing) of extreme data.
Extreme data is defined as data that exhibits one or more of the following characteristics, to an extent that makes current technologies fail: increasing volume, speed, variety; complexity/diversity/multilinguality of data; the dispersed data sources; sparse/missing/insufficient data/extreme variations in values).
The technologies and solutions are expected to discover and distil meaningful, reliable and useful data from heterogeneous and dispersed/scarce sources and deliver it to the requesting application/user with minimal delay and in the appropriate format.
Extreme Near-Data Processing Platform
Why Locality ?
Volume
Privacy
Low latency
Hardware Acceleration
Objectives
The main goal is to design an Extreme near-data processing platform to enable consumption, mining and processing of distributed and federated data without needing to master the logistics of data access across heterogeneous data locations and pools.
KPIs
Software components
Workpackages
WP2 – Global Architecture
Objectives:
• Design the overall system architecture of the NEARDATA software (T2.1).
• Provide and implement a set of interfaces and APIs to integrate the different components and software prototypes of the platform (T2.2).
• Describe stress-test scenarios and benchmarking framework (T2.3) and validate them in International Data Spaces (T2.4).
WP3 – Data Plane: Extreme Data Connectors
Objectives:
• Development of Serverless Data Connector platform for Data connectors (T3.1).
• Deploy Stream data connectors as stream operators (T3.2, T3.3).
• Validate High Performance data connectors using Hardware acceleration (T3.3).
• Integration with popular analytic platforms (T3.5).
Data Connectors
Results:
• Lithops: Resource Auto-Scaling.
• Pravega: Throughput and data speed improvements (x1,06 faster than Kafka and x1,4 faster than Pulsar).
• SCONE: Confidential Computing.
• Data Connectors:
WP4 – Control Plane: Confidential Data Orchestration
Objectives:
• Develop a Data broker component providing data governance and data orchestration (T4.1).
• Implement confidential data compute and exchange mechanisms leveraging TEEs (T4.2, T4.3).
• Develop confidential data orchestration mechanisms including federated learning (T4.4, T4.5).
Results:
• Integration of Confidential Compute Layer of Data Broker with Lithops and use-cases.
AI Component
WP5 – Extreme Health Use Cases
Objectives:
• Optimize Use Case workloads using machine learning techniques that leverage WP3 and WP4 technical achievements (T5.1).
• Validate the platform in Genomics, Metabolomics, and Surgery use cases with complex pipelines involving data connectors (T5.2, T5.3, T5.4, T5.5, T5.6).
• Create and validate open libraries of data connectors in the different use cases (T5.2, T5.3, T5.4, T5.5, T5.6).
WP5 – Extreme Health Use Cases
Metabolomics Use-case.
WP5 – Extreme Health Use Cases
Genomics Use-cases.
WP5 – Extreme Health Use Cases
Surgery Use-case.
WP6 – Promoting impact
Objectives:
• Inform stakeholders (such as industry, scientific communities, EU officers, commission and administration; general public and media) about the progress of the project. A special attention will be paid to reach SMEs in involvement and utilization of the results.
• Encourage the production of articles, reports and demonstrations of the project results (T6.1).
• Perform relevant collaboration and exploitation activities (T6.2).
• Monitor standard setting bodies activities and contribute our achievements (T6.3).
Software Outcomes
Agenda
Thank you