Advances in High Performance Computing and Deep Learning: Data Engineering and Data Science
Big Data Systems and HPC
1
The IEEE DSS-2020 Virtual Conference
http://www.ieee-cybermatics.org/2020/dss/
Digital Science Center
Big Data Systems and HPC
Abstract
2
Digital Science Center
Big Data Systems and HPC
What we will discuss today
3
Digital Science Center
Big Data Systems and HPC
Data Engineering and Deep Learning
What do we need to support?
4
12/7/2019
Digital Science Center
Big Data Systems and HPC
Deep Learning Workflow
5
Workflow often divide into two:
Data => Information preprocessing -- Hadoop, Spark, Twister2, Scikit-Learn
Information => Knowledge Compute intensive step Cylon enhanced Spark Twister2, PyTorch and Tensorflow
Post-Processing
Data
Digital Science Center
Big Data Systems and HPC
6
ML Code
NIPS 2015 http://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-systems.pdf
This well-known paper points out that parallel high-performance machine learning is perhaps most fun but just a part of system. We need to integrate in the other data and orchestration components.
This integration is not very good or easy partly because data management systems like Spark are JVM-based which doesn’t cleanly link to C++, Python world of high-performance ML
ML code module is itself built up hierarchically from Numpy and Pandas operations (if Python)
Need to assemble 10 large modules into full workflow and efficiently execute Numpy/Pandas inside modules
Integrating Data Engineering and Data Science
Digital Science Center
Big Data Systems and HPC
Some Important Trends I
7
Digital Science Center
Big Data Systems and HPC
Some Important Trends II
8
Digital Science Center
Big Data Systems and HPC
Data Engineering and Deep Learning
A solution with Cylon and Twister2
9
12/7/2019
Digital Science Center
Big Data Systems and HPC
Digital Science Center
Big Data Systems and HPC
Data Engineering
Digital Science Center
Big Data Systems and HPC
Twister2
Big Data Processing�Eco-System
Dataflow
API
Dataflow
API
Twister2
Linear/Relational Algebra Operators
Distributed Linear/Relational Algebra Operators [C++]
Distributed Relational Communication Operations [C++]
Communication Kernels
Twister2 one of 5 possible engines for Apache Beam, which implements a rich data processing (dataflow) workflow environment – other engines are Spark, Flink, Samza, Google Cloud Dataflow
Cylon
Lines of Code�Twister2 125000 (Java, Python)
Cylon 15000 (C++, Python, Java)
Digital Science Center
Big Data Systems and HPC
244 Pandas Operators: Core Data Functions
pandas.DataFrame, index, columns, dtypes, info, select_dtypes, values, axes, ndim, size, shape, memory_usage, empty, astype, convert_dtypes, infer_objects, copy, bool, head, at, iat, loc, iloc, insert, __iter__, items, iteritems, keys, iterrows, itertuples, lookup, pop, tail, xs, get, isin, where, mask, query, add, sub, mul, div, truediv, floordiv, mod, pow, dot, radd, rsub, rmul, rdiv, rtruediv, rfloordiv, rmod, rpow, lt, gt, le, ge, ne, eq, combine, combine_first, apply, applymap, pipe, agg, aggregate, transform, groupby, rolling, expanding, ewm, abs, all, any, clip, corr, corrwith, count, cov, cummax, cummin, cumprod, cumsum, describe, diff, eval, kurt, kurtosis, mad, max, mean, median, min, mode, pct_change, prod, product, quantile, rank, round, sem, skew, sum, std, var, nunique, value_counts, add_prefix, add_suffix, align, at_time, between_time, drop, drop_duplicates, duplicated, equals, filter, first, head, idxmax, idxmin, last, reindex, reindex_like, rename, rename_axis, reset_index, sample, set_axis, set_index, tail, take, truncate, backfill, bfill, dropna, ffill, fillna, interpolate, isna, isnull, notna, notnull, pad, replace, droplevel, pivot, pivot_table, reorder_levels, sort_values, sort_index, nlargest, nsmallest, swaplevel, stack, unstack, swapaxes, melt, explode, squeeze, to_xarray, T, transpose, append, assign, compare, join, merge, update, asfreq, asof, shift, slice_shift, tshift, first_valid_index, last_valid_index, resample, to_period, to_timestamp, tz_convert, tz_localize, attrs, plot, plot.area, plot.bar, plot.barh, plot.box, plot.density, plot.hexbin, plot.hist, plot.kde, plot.line, plot.pie, plot.scatter, boxplot, hist, sparse.density, sparse.from_spmatrix, sparse.to_coo, sparse.to_dense, from_dict, from_records, to_parquet, to_pickle, to_csv, to_hdf, to_sql, to_dict, to_excel, to_json, to_html, to_feather, to_latex, to_stata, to_gbq, to_records, to_string, to_clipboard, to_markdown, style,
13
Digital Science Center
Big Data Systems and HPC
Two Ecosystems
Enterprise: Java
Pre and Post Data Engineering
Research Labs, Universities: Python
Central Data Science (Deep Learning)
GOAL: High Performance in each ecosystem �and high-performance integration between ecosystems!
Digital Science Center
Big Data Systems and HPC
Twister2 Benchmarks with Big Data Processing Ecosystem
This benchmarks full (micro)services
Digital Science Center
Big Data Systems and HPC
Cylon: A High Performance Distributed Data Table
16
Digital Science Center
Big Data Systems and HPC
Cylon Architecture
Builds on Apache Arrow to��Link Python world: Jupyter, Numpy, Pandas, Modin
(Cython links Python to C++ Cylon)
with Java world�Spark, Twister2
and C++/CUDA world�High performance deep learning on PyTorch and Tensorflow
Table/Dataframe interface for usability
Digital Science Center
Big Data Systems and HPC
Strong Scaling Comparison with Other Frameworks
Inner join
Union
Digital Science Center
Big Data Systems and HPC
Large Scale Experiments with PySpark and PyCylon
Digital Science Center
Big Data Systems and HPC
Strong Scaling Comparison with Other Frameworks
Aggregations
Group-by + aggregations
Digital Science Center
Big Data Systems and HPC
Cylon Performance with Language Bindings
Digital Science Center
Big Data Systems and HPC
Interoperability Goals for Cylon
Digital Science Center
Big Data Systems and HPC
Dimension Reduction from Classic MDS (improved by DA) to Deep Learning
Example of DL replacing previous methods
23
12/7/2019
Digital Science Center
Big Data Systems and HPC
MDS Approaches: Deep Learning v Previous Best Practice
1500-->128-->3-->128-->1500
Digital Science Center
Big Data Systems and HPC
Silhouette Coefficient SC
Silhouette Coefficient for different approaches and differing K
Method | Silhouette Coefficient (10 random sample runs) | ||
Min | Max(Best) | Average | |
DAMDS | NA | -0.2199 | -0.2199 |
One Hot Encoding | NA | -0.5022 | -0.5022 |
References (K = 25) | -0.388 | -0.233 | -0.309 |
References (K = 50) | -0.397 | -0.229 | -0.289 |
References (K = 100) | -0.371 | -0.238 | -0.278 |
References (K = 200) | -0.331 | -0.236 | -0.268 |
References (K = 400) | -0.306 | -0.231 | -0.262 |
References (K = 800) | -0.300 | -0.219 | -0.256 |
References (K = 1000) | -0.308 | -0.209 | -0.257 |
References (K = 1500) | -0.311 | -0.193 | -0.248 |
References (K = 2000) | -0.297 | -0.207 | -0.264 |
References (K = 4000) | -0.357 | -0.294 | -0.322 |
K=6000 larger network: 6000x512x128x32x3 | -0.174 | -0.238 | |
a(i) - sum of distance to points in same cluster
b(i) - min sum of distance to points in other clusters
Digital Science Center
Big Data Systems and HPC
170K fungi Gene sequence HeatMaps
Heatmaps of Smith-Waterman distance vs projected distance in 3D
DAMDS vs Smith-Waterman
OHE vs Smith-Waterman
100 References vs Smith-Waterman
1K References vs Smith-Waterman
(a)
(b)
(c)
1.5K References vs Smith-Waterman
(a)
(b)
(c)
6K References�Larger Network
6000x512x128x32x3
(d)
(d)
Digital Science Center
Big Data Systems and HPC
Out of Sample Data
Some Lessons
Out Of Sample Data Points | Incorrect Classification | Accuracy |
2000 | 4/1000 | 99.8% |
4000 | 17/4000 | 99.57% |
8000 | 17/8000 | 99.78% |
Digital Science Center
Big Data Systems and HPC
Deep Learning Surrogates
28
12/7/2019
Digital Science Center
Big Data Systems and HPC
Some Important Trends II
29
Digital Science Center
Big Data Systems and HPC
Let’s look at ML for HPC
30
Digital Science Center
Big Data Systems and HPC
Up to two billion times acceleration of scientific simulations with deep neural architecture search
31
Digital Science Center
Big Data Systems and HPC
Examples of ML for HPC (work with JCS Kadupitiya, Vikram Jadhao)
32
→ 106 as Nlookup → ∞
The Learning Net
Direct simulation compared to Surrogates
Digital Science Center
Big Data Systems and HPC
INSILICO MEDICINE USED CREATIVE AI TO DESIGN POTENTIAL DRUGS IN JUST 21 DAYS
33
1/7/2020
September 4 2019 News Item
Digital Science Center
Big Data Systems and HPC
Operator Formulation of Deep Learning Inference
34
Digital Science Center
Big Data Systems and HPC
Learn Newton’s laws with Recurrent Neural Networks
35
are 4000 times longer � and also learn potential.
RNN Error2 up to step size dT=4 and total time 106
Verlet error2�dT = 0.01, 0.1
10-5
1023
101
JCS Kadupitiya, Vikram Jadhao
Digital Science Center
Big Data Systems and HPC
Status of Surrogates and ML for HPC
36
Digital Science Center
Big Data Systems and HPC
Futures of Surrogates and Computational Science
37
Digital Science Center
Big Data Systems and HPC
Research Issues
Neural Search
Uncertainty Quantification
Minimizing Size of Needed Training Set
38
Digital Science Center
Big Data Systems and HPC
Diversion on Time Series
39
12/7/2019
Digital Science Center
Big Data Systems and HPC
Times Series represented by Deep Learning
40
Digital Science Center
Big Data Systems and HPC
Basic Spatial (bag) Time Series
41
Forecast the Future �(any number of time units
any number properties)
Predict Now
or Seq2Seq map
as in
English to French
or
rainfall to runoff
Input Properties
Static e.g. %Seniors
Dynamic e.g. Covid cases per day
Space x
(Different data sources, not necessarily nearby)
Forecast the Future
Time t
Seq 2 Seq
Data Analysis Unit
Time sequence at one space point
For Natural Language Processing, space points are different paragraphs or books. A few sentences at each point.
Digital Science Center
Big Data Systems and HPC
General Deep Learning Strategies II
42
Hybrid Transformer (for encoder) and LSTM (for decoder)
Pure LSTM
Two different models used
Merge
Final
Initial
LSTM Layer
LSTM Layer
Outputs
Final
Initial
LSTM Layer
LSTM Layer
Outputs
Inputs
optional but best results if you do this
Digital Science Center
Big Data Systems and HPC
Search Strategies in Hybrid (Science) Transformer
43
Each color is a different location with 5 time values�N=10 W=5 in exampler
Digital Science Center
Big Data Systems and HPC
44
Time Series
Benchmarks
We are looking at examples of time series including hydrology (rainfall) and epidemiology as in COVID infection and fatality data.
Latter dataset is currently 314 cities and 205 days
Models are LSTM and Transformer�W=9 N=159
Digital Science Center
Big Data Systems and HPC
Hybrid Transformer (Attention+LSTM) with intrinsic error for 314 cities/counties and 159 days
45
Digital Science Center
Big Data Systems and HPC
CAMELS Hydrology Data Sequence Length 21. Simple search along axes. 671 locations, 7000 days
46
Runoff with some locations missing data so location sum observation not plotted - only ~2000 missing points over 671 locations, 7000 days
Digital Science Center
Big Data Systems and HPC
Implementation Issues
47
Digital Science Center
Big Data Systems and HPC
Collection of Time Series Machine Learning Algorithms (MLPerf)
48
Areas | Applications | Model | Data sets | Papers |
Cars, Taxis, Freeway Detectors | TT-RNN , BNN, LSTM | Taxi/Uber trips [2-5] | [6-8] | |
Wearables, Medical instruments: EEG, ECG, ERP, Patient Data | LSTM, RNN | OPPORTUNITY [9-10], | [16-20] | |
Intrusion, classify traffic, anomaly detection | LSTM | GPL loop dataset [21], SherLock [22] | [21, 23-25] | |
Household electric use, Economic, Finance, Demographics, Industry | CNN, RNN | Household electric [26], M4Competition [27], | [28-29] | |
Stock Prices versus time | CNN, RNN | Available academically from Wharton [30] | [31] | |
Climate, Tokamak | Markov, RNN | USHCN climate [32] | [33-35] | |
Events | LSTM | Enterprise SW system [36] | [36-37] | |
Language and Translation | Pre-trained Data | Transformer [38] | [39-40] | [41-42]Mesh Tensorflow |
All-Neural On-Device Speech Recognizer | RNN-T | | [43] | |
IndyCar Racing | Real-time car and track detectors | HTM, LSTM | | [44] |
Online Clustering | Available from Twitter | [45-46] |
Xinyuan Huang from Cisco and MLPerf
Digital Science Center
Big Data Systems and HPC
Conclusions
49
12/7/2019
Digital Science Center
Big Data Systems and HPC
MLCommons (MLPerf) Consortium Deep Learning Benchmarks
Some Relevant Working Groups
50
Benchmarking, Datasets, Best Practices
Total Effort ~50 FTE
Digital Science Center
Big Data Systems and HPC
Conclusions
51
Digital Science Center
Big Data Systems and HPC
Call for Action!
52
Digital Science Center
Big Data Systems and HPC
Thank you!
External Collaborators at Argonne National Lab, Arizona State University, Kansas, Rutgers, Stony Brook, UT Knoxville, University of Virginia and MLPerf
Indiana University Digital Science Center:
Faculty: David Crandall, James Glazier, Vikram Jadhao, Judy Qiu and others
Staff: Josh Ballard, Gary Miksik, Fugang Wang, Chathura Widanage,
Researchers: Gurhan Gunduz, Supun Kamburugamuve, Ahmet Uyar, Gregor von Laszewski
Students: Vibhatha Abeykoon, Bo Feng, JCS Kadupitiya, Niranda Perera, Pulasthi Wickramasinghe,
53
Digital Science Center
Big Data Systems and HPC