1 of 15

ADVANCED VISUALIZATION OF POWER, TEMPERATURE, AND ENERGY METRICS IN HPE CRAYEX SYSTEMS

Sustainable HPC State of the Practice Workshop 2024

Authors:

Lavanya L

Stefan Ceballos

September 24, 2024

  • 1

© 2024 Hewlett Packard Enterprise Development LP

2 of 15

Agenda

  • 2

Introduction to HPCM & Monitoring Framework

Limitations with the traditional approach to monitoring

HPE Clusterview Plugin

Automation of Clusterview Configuration

Challenges with Clusterview Configuration

Visual Clusterview Dashboard Overview

© 2024 Hewlett Packard Enterprise Development LP

Future Scope

3 of 15

HPE Performance Cluster Manager

All you need to manage your HPE clusters, keep them healthy, and running at peak performance

  • 3

Flexible, easy to use system management solution offering system administrators all the tools they need to turn even the most complex hardware into easily manageable systems capable of accommodating a growing variety of workloads.

Powerful

Comprehensive set of tools you need to manage all aspects of your cluster.

Productive

Designed to maximize productivity of your cluster, automate actions, and optimize running costs.

Scalable

Manage systems from dozens of nodes up to Exascale.

 

Anywhere

Suitable for both on-premises as well as hybrid HPC deployments.

Flexible

Customize monitoring, alerts, and management actions to best suit your needs.

Proven

Used by hundreds of customers around the globe.

© 2024 Hewlett Packard Enterprise Development LP

4 of 15

Monitoring Framework

  • 4

Data Sources

HPE Cray EX, HPE Cray XD and HPE Apollo

Redfish End points & Sensor Data

Fabric—Slingshot & InfiniBand

Events and Telemetry

WLM—PBS Pro & SLURM

Events and Telemetry

Compute systems

CPU, GPU, Memory, Disk Telemetry and Events

CDU Events and Telemetry

Storage and Filesystems

GlusterFS

Power and Heartbeat Monitoring

Syslog

Data Pipeline

Persistent Storage

Monitoring Framework

Victoria Metrics

Service Infrastructure Monitoring

OpenSearch

Victoria Metrics

Mosquitto

subsmon

telegraf

pcim

Kafka

User Interfaces

Grafana

OpenSearch

Alertmanager

Dashboards

Alerts

Rest API

AIOps

Management hardware and software

Customer Integration

© 2024 Hewlett Packard Enterprise Development LP

5 of 15

Limitations with the traditional approach to monitoring

  • 5
  • Lack of Granularity: Standard dashboards often fail to provide detailed insights at a fine level (e.g., per-node or component-wise data), making it difficult to pinpoint issues.

  • Overwhelming Volume of Data: Exascale and pre-exascale systems generate massive amounts of telemetry data (temperature, power metrics, job status, etc.), leading to information overload in traditional dashboards.

  • Limited Scalability: Standard dashboards may struggle to scale up effectively for systems with tens of thousands of nodes, leading to performance bottlenecks or missing information.

  • Complexity in Monitoring Failures: Monitoring co-located failures in large clusters becomes difficult, as native panels may not show correlated data effectively.

  • Inadequate Node Status Tracking: Real-time tracking of node status during system bring-up (e.g., validation tests) is cumbersome without a focused view.

© 2024 Hewlett Packard Enterprise Development LP

6 of 15

Challenges with the traditional approach to monitoring

  • 66

© 2024 Hewlett Packard Enterprise Development LP

7 of 15

Challenges with the traditional approach to monitoring

  • 7

© 2024 Hewlett Packard Enterprise Development LP

8 of 15

HPE Clusterview Plugin

  • 8

KEY FEATURES OF THE HPE CLUSTERVIEW PLUGIN

Hierarchical Data Grouping

Data is organized hierarchically, with each layer of the hierarchy (e.g., chassis, blades, nodes) to show detailed views.

Regex-based Data Parsing

It uses regular expressions (regex) to extract and format the data.

Custom Layouts

It supports various display layouts (e.g., horizontal, vertical, flow, grid), allowing for more complex arrangements that better reflect the structure of an HPC system.

Interactive and Clickable Nodes

Each node can be clickable with customizable URLs, allowing for a more interactive user experience.

The HPE Clusterview plugin for Grafana is designed specifically for high-performance computing (HPC) environments, offering a highly detailed view of various data points in a dense, hierarchical format.

It focuses on spatial representation and the operational hierarchy of the HPC infrastructure

It makes it ideal for large-scale systems where managing and visualizing complex data across multiple nodes is critical.

© 2024 Hewlett Packard Enterprise Development LP

9 of 15

Visualizing Slingshot Metrics

© 2024 Hewlett Packard Enterprise Development LP

  • 9

Switches are displayed in group orientation

Green represents nominal, red represents anomalous

HPE Clusterview Plugin

© 2024 Hewlett Packard Enterprise Development LP

10 of 15

Visualizing Slingshot Metrics

  • 10

Details of ports inside of a switch

Green represents normal�Red represents anomalous�Grey represents no data

HPE Clusterview Plugin

© 2024 Hewlett Packard Enterprise Development LP

11 of 15

Visualizing Heartbeat of Aurora Compute Nodes

© 2024 Hewlett Packard Enterprise Development LP

  • 11

HPE Clusterview Plugin

Nodes are displayed in rack orientation

© 2024 Hewlett Packard Enterprise Development LP

12 of 15

Challenges with Clusterview Configuration

  • 12
  • Manual Configuration Complexity: Configuring the Clusterview plugin involves �manually writing regex patterns for each group (rack, cabinet, chassis, slot, blade, node) �based on the cluster’s hardware.

  • Customization for Unique Hardware: Each cluster's configuration is unique, meaning �that regex patterns cannot be duplicated.

  • Scaling Issues for Large Systems: Large systems, like the Frontier exascale machine with �over 9,400 nodes, demand extensive manual configuration efforts, leading to inefficiencies �and potential delays in system bring-up.

  • Error-Prone Layout Configuration: Mistakes in the regex setup can lead to incorrect �representations of the system's physical layout. This misrepresentation can cause delays in �identifying genuine issues, particularly in environments with thousands of nodes, such as �exascale systems.

© 2024 Hewlett Packard Enterprise Development LP

13 of 15

Automation of Clusterview Configuration

  • 13

The automated solution addresses the issues by detecting the hardware type and location within the data center, enabling the real-time configuration of dashboards.

The key components of the automation of Clusterview configuration:

  • Command Line Interface: Users can execute the automation process with specific requirements using the command line tool. This targeted approach ensures that only the necessary dashboards are updated, providing flexibility and improving efficiency.

  • Data Collection: The framework collects metrics, identifying hardware types and configurations. Detailed data is gathered for nodes, racks, cabinets, chassis, slots, blades, and more, ensuring that the data presented in dashboards is accurate and meaningful.

  • Data Grouping: Once data is collected, the framework dynamically generates regular expressions (regex) to group nodes at different levels. It automatically creates visual elements like borders and labels, providing a clear and structured view of the node hierarchy.

© 2024 Hewlett Packard Enterprise Development LP

14 of 15

Future Scope

  • 14

Enhanced Visual Scaling

Plans to auto-scale node representations based on cluster size and screen orientation, improving readability and aesthetics across different devices.

Integration with Other HPE Solutions

Porting and integrating the solution with future versions of HPE cluster management software like Cray System Management (CSM) and System Monitoring Health.

Automation for Additional Components

The automation framework aims to extend beyond node layouts to cover additional cluster components, including network fabrics, CDUs, and PDUs, allowing for a more comprehensive and automated cluster configuration process.

Open-Source Offering

The goal is to make the automation framework available as an open-source tool, enhancing accessibility alongside the publicly available HPE Clusterview plugin.

© 2024 Hewlett Packard Enterprise Development LP

15 of 15

Thank You!

Email: lavanya.l@hpe.com