Unified System- and Network-Level Monitoring for AI-Driven Anomaly Detection on Chameleon Cloud
Arin Rahman
arahman6@miners.utep.edu
The University of Texas at El Paso
Project Overview
Feature Ranking using Gain Ratio and Feature Subset generation
top 10%
top 12%
top 14%
…
top 20%
Gain Ratio
F-1 Score vs PAPI + LDMS feature subset
Best Model Performance Comparison Across Feature Sets
Infrastructure Requirements and Usage
Challenges and Insights
Services, Tools, and Workflow Support
Recommendations for AI Research Testbeds
Hardware Visibility: Testbeds must provide persistent access to low-level hardware performance counters.
Native Orchestration: Providing container-native templates (like Docker Swarm or Kubernetes) simplifies complex security experiments.
Integrated Pipelines: A need exists for "telemetry-to-ML" pipelines to bridge the gap between raw data collection and AI evaluation.
Publications
1.Rahman, A., Moore, S. V., & Tosh, D. K. (2026). Design of a unified monitoring tool for detecting anomalies in high performance computing systems. In 2026 IEEE 23rd Consumer Communications & Networking Conference (CCNC) (pp. 1–2). IEEE.
2. Rahman, A., Moore, S. V., & Tosh, D. K. Unified System- and Network-Level Monitoring for AI-Driven Anomaly Detection on Chameleon Cloud. Manuscript under review at IEEE ClusterCom.
Acknowledgement
This research was supported by the National Science Foundation (NSF) under award #2346423.
Questions?
Email: arahman6@miners.utep.edu