1 of 1

Machine learning-based application grouping and knowledge extraction for HPC I/O

Scientific Achievement

Developed an approach for structuring HPC jobs in a traversable hierarchy,

that is in turn used for efficient knowledge extraction of I/O throughput issues. This framework is available as an opensource tool (Gauge) for HPC system workload exploration.

Significance and Impact

Diagnosing storage system utilization and application I/O patterns is slow and does not scale. Our approach, through a combination of feature engineering, hierarchical clustering and interpretable models provides a holistic framework to quickly analyze groups of similar jobs and extract application I/O insights at different granularity of groupings.

Research Details

    • Hierarchical clustering to create an easy-to-navigate tree of jobs
    • ML models predicting I/O throughput trained on clusters of HPC jobs reveal local insight into sources of I/O throughput variance
    • New training set generation method tackles extrapolation gap
    • Demonstrated on ALCF Theta system

Fig.2: Visualization of a cluster of jobs, where the blue application significantly outperforms the red one. Gauge helps diagnose these I/O problems.

Fig.1: ALCF Theta Workload Hierarchy

ANL: S. Madireddy, P. Balaprakash, P. Carns, R. Ross

TAMU: M.Kinsy, M. Isakov, E. del Rodario

M. Isakov, E. del Rosario, S. Madireddy, P. Balaprakash, P. Carns, R. Ross, and M. Kinsy, “HPC I/O throughput bottleneck analysis with explainable local models,”

in SC’20: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2020.