Machine learning-based application grouping and knowledge extraction for HPC I/O
Scientific Achievement
Developed an approach for structuring HPC jobs in a traversable hierarchy,
that is in turn used for efficient knowledge extraction of I/O throughput issues. This framework is available as an opensource tool (Gauge) for HPC system workload exploration.
Significance and Impact
Diagnosing storage system utilization and application I/O patterns is slow and does not scale. Our approach, through a combination of feature engineering, hierarchical clustering and interpretable models provides a holistic framework to quickly analyze groups of similar jobs and extract application I/O insights at different granularity of groupings.
Research Details
Hierarchical clustering to create an easy-to-navigate tree of jobs
ML models predicting I/O throughput trained on clusters of HPC jobs reveal local insight into sources of I/O throughput variance
New training set generation method tackles extrapolation gap
Demonstrated on ALCF Theta system
Fig.2: Visualization of a cluster of jobs, where the blue application significantly outperforms the red one. Gauge helps diagnose these I/O problems.
Fig.1: ALCF Theta Workload Hierarchy
ANL: S. Madireddy, P. Balaprakash, P. Carns, R. Ross
TAMU: M.Kinsy, M. Isakov, E. del Rodario
M. Isakov, E. del Rosario, S. Madireddy, P. Balaprakash, P. Carns, R. Ross, and M. Kinsy, “HPC I/O throughput bottleneck analysis with explainable local models,”
in SC’20: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2020.