Farm impact of CDCE cooling failure on 6/15 (Will, Chris, Ricky, Kevin)
On-board automated temperature monitoring & shutdown script functioned as designed, and powered down hundreds of machines once ambient or inlet temperature reached 100F
Affected shared pool nodes (spool0XYZ) and ATLAS T1 nodes in CDCE.
~6 machines could not be recovered after the incident
All but one were older (~9 years old) hosts
STAR CME Analysis Activity
DOE priority to complete this analysis in August
All STAR nodes added to the shared pool
NPP management requested ~40% of other experiments’ slots in shared pool moved to CME group in the short term
CVMFS now functional with REANA on our k8s cluster (Chris)
Code changes to the reana-job-controller image to allow use of CVMFS mounts
HPSS Upgrade (Tim, Iris, Will, Chris)
We are on schedule for previously announced 8/2 - 8/5 HPSS 8.3.10 upgrade
Will send out reminder email next week
Successfully tested our modified HPSS batch and Lustre Copytool code against 8.3.10
PFTP will be available for testing next week
GPFS client upgrade
Upgrade required in the near future - requires draining jobs
Does not impact ATLAS T1 as GPFS is not mounted on their nodes
Can this be done all at once, or do experiments prefer in a rolling (⅓ of the farm affected at a time) upgrade?