N-DISE: NDN for Data Intensive Science
Experiments
Edmund Yeh
Large-Scale Data-Intensive Science
LHC: Large Hadron Collider
LSST: Large Synoptic Survey Telescope
SKA: Square Kilometer Array
Genomics
Challenges of Data-Intensive Science�
- LHC high energy physics, genomics, LSST, SKA, LIGO, EHT
- Indexing, security, storage, distribution, analysis, learning
- Coordinated use of computing, storage, network resources
- Incremental solutions; developed in isolation; replicated efforts
- Current computer networks/systems focus on addresses,
processes, servers, connections
- Applications care about data
Closing the Gap: Data-Centric Networking�
Data-Centric Networking�
- Data production: naming, authenticating, securing data directly
- Delivering data using names enables scalable data retrieval
- in-network caching
- automated joint caching and forwarding
- multicast delivery
NSF N-DISE and SANDIE Projects
N-DISE/SANDIE Results and Demos
N-DISE WAN testbed
Throughput Test
Starlight
Caltech
SC22 booth
Consumer
& Forwarder
Consumer & Forwarder
3610
1875
Forwarder & Producer
3611
Throughput Test: Local
Throughput Test
Throughput Test
VIP caching Test
VIP caching Test
VIP caching Test
VIP caching Test
XRootD and the Open Storage System Plugin
- is a software framework used for file management at CERN
- enables developers to implement various types of plugins (e.g., filesystem, caching, security)
- is written in C++, though there is ongoing work for a Golang version�
- is a C++ dynamic library that offers support for all related POSIX file system calls (e.g., open, close, read, opendir, readdir)
- bridges IO to/from Ceph File System without using a kernel
NDN-based XRootD OSS Plugin
/ndnc/xrootd/store/mc/file1/32=metadata for open and fstat
/ndnc/xrootd/store/mc/file1/v1/seg=3 for read
File Transfer Using NDN-based XRootD Plugin
NDNc File Transfer Client vs. XRootD OSS Plugin
- Uses two threads: one for encoding interests, other for reading and decoding data
- The two threads operate asynchronously and do not communicate
- This approach eliminates overhead and benefits from parallelization, which leads to improved throughput
- CMS configuration only allows synchronous requests
- Each worker thread will read 2MB of data at an offset from a file and will wait for the response before continuing, which limits throughput
FPGA Acceleration of NDN-DPDK Forwarder
Naming Compute and Data
NDN Data Lake
NDN-TR70 - Utilizing NDN DPDK for Kubernetes Genomics Data Lake
Sankalpa Timilsina, Justin Presley, David Reddick, Susmit Shannigrahi, Tennessee Tech,�Xusheng Ai, Coleman Mcknight, Alex Feltus�
Compute Placement based on Names
Compute Placement based on Names
N-DISE Demo at SC22
N-DISE team:
Northeastern: E. Yeh, Y. Wu, V. Mutlu, Y. Liu
Caltech: H. Newman, C. Iordache, R. Sirvinskas, J. Balcas
UCLA: L. Zhang, J. Cong, S. Song, M. Lo
Tennesee Tech: S. Shannigrahi, S. Timilsina
NIST: D. Pesavento, J. Shi, L. Benmohamed
Special thanks to:
ESnet: T. Lehman
CENIC: S. Bellamine
StarLight: J. Mambretti, F. Yeh, J. Chen, S. Yu
Internet2: M. Zekauskas
RNP: M. Schwarz
N-DISE Long-Term Goals
Thank you!