1 of 1

IT Fabric (6/24)

  • Farm impact of CDCE cooling failure on 6/15 (Will, Chris, Ricky, Kevin)
    • On-board automated temperature monitoring & shutdown script functioned as designed, and powered down hundreds of machines once ambient or inlet temperature reached 100F
    • Affected shared pool nodes (spool0XYZ) and ATLAS T1 nodes in CDCE.
    • ~6 machines could not be recovered after the incident
      • All but one were older (~9 years old) hosts

  • STAR CME Analysis Activity
    • DOE priority to complete this analysis in August
    • All STAR nodes added to the shared pool
    • NPP management requested ~40% of other experiments’ slots in shared pool moved to CME group in the short term

  • CVMFS now functional with REANA on our k8s cluster (Chris)
    • Code changes to the reana-job-controller image to allow use of CVMFS mounts

  • HPSS Upgrade (Tim, Iris, Will, Chris)
    • We are on schedule for previously announced 8/2 - 8/5 HPSS 8.3.10 upgrade
      • Will send out reminder email next week
    • Successfully tested our modified HPSS batch and Lustre Copytool code against 8.3.10
    • PFTP will be available for testing next week

  • GPFS client upgrade
    • Upgrade required in the near future - requires draining jobs
      • Does not impact ATLAS T1 as GPFS is not mounted on their nodes
    • Can this be done all at once, or do experiments prefer in a rolling (⅓ of the farm affected at a time) upgrade?
    • Do we need to delay due to CME activity?