1 of 20

AGLT2 Pre-scrubbing Info

Shawn McKee/University of Michigan

Dan Hayden, Philippe Laurens, Wenjing Wu

Pre-scrubbing at BNL

https://indico.bnl.gov/event/16189/ June 28, 2022

2 of 20

AGLT2 Overview and Stats

2

AGLT2 Site Report

  • HTCondor Cluster:
    • 325 Worker Nodes, 15.1 kCores, 175 kHS06, Avg. 11.59 HS06/core
    • 2GB~6.3GB RAM/core, 1000 job slots for High Memory Queue (6GB/core)
    • 14GB~52GB Disk/core, supports Merge Queue with higher disk requirement.
    • 2x 10or25Gpbs bonded NICs for 90% work nodes (2x 1Gpbs for the 10% oldest nodes)
  • dCache storage:
    • 12PB deployed (5.3 PB@MSU, 6.7PB@UM), 11 PB in space tokens
    • 2x 25 or 4x 10Gpbs bonded NICs for almost all storage nodes (2x10Gbps for 5x oldest)
  • Continue to use and optimize the ATLAS@home backfilling on the HTCondor cluster.
    • optimize the cgroup and BOINC configurations to reduce the impacts of the backfilling on the CPU Efficiency of the Condor jobs.
  • Networking (resilient 100 Gbps):
    • We have completely (and very successfully) replaced our UM LAN and UM/MSU WAN network equipment starting last summer and finishing up WAN and routing changes in March 2022. Converted to all optical and retired our DAC (copper) high-speed (25G, 40G, 100G) cables.
    • MSU migrated to a new data center, which provided new network devices, optics, cabling at no cost to AGLT2

3 of 20

WBS 2.3.2.1 AGLT2 Personnel / Effort

3

Name

Effort (FTE)

Wenjing (Wendy) Dronen (UM)

1.0

Philippe Laurens (MSU)

0.95 (temp; normally 0.85)

Shawn McKee (UM)

0.1

Dan Hayden (MSU)

0.05

Mike Nila (MSU)

0.0 (retired)

Total

2.1

4 of 20

AGLT2 Equipment

4

End of 2021 and early 2022 purchases already installed:

  • 18x R6525 (AMD 7302 32HT, 1152 cores @16.42 HS06/core, Total 18.9 kHS06)
  • 8x R740xd2 (12T drives, total 1888 TB usable in dCache)
  • 3x VMware hosts, 1x NVMe storage for VMware, 1x NVMe node for SLATE, 1x Zeek node
  • 8x R740xd2 (18T drives, total 2880 TB usable in dCache)

But still waiting on (estimated delivery July 2022)

  • 29x R6525 (AMD 7413 48HT, 2784 cores, HS06 TBD, Total prelim est. over 45 kHS06)

Networking and site upgrades

  • MSU equipment relocated to campus data center. 12x33kW racks. Dual, true redundant power. �Multipath 100G WAN. Dual/redundant data switches 25G/port in each rack. Room for expansion.
  • UM networking gear and cabling upgraded. Multipath 100G WAN [is that a true and compact/sufficient summary, implies the UM-MSU failover]
  • Separate multipath 100G MSU-UM inter-site connectivity.
  • Storage Nodes: 2*25Gpbs, Work Nodes: 2*10Gpbs, VMware Cluster: 4*25Gpbs

​

​

5 of 20

Software and Technology Plans

5

AGLT2 runs a number of software packages required for an ATLAS site:

  • OSG 3.6
  • HTCondor 9.0.13
  • HTCondor-CE 5.1.5
  • dCache 7.2.19

Currently our base operating system is CentOS 7.9 but we would like to migrate to either RHEL9 or a RHEL9 compatible OS (Rocky 9, Almalinux 9, CentOS Streams 9).

To host and manage critical services we rely upon VMware, which provide high availability and supports live migration of services to allow hardware, firmware and software updates

  • VMware 6.7U3 (plan to upgrade to 7.x before EOL in September)

Storage: we also run Lustre, NFS servers, and AFS cell and have collaborative access to Ceph.

We have a combination of custom built monitoring tools, along with CheckMK, Elasticsearch, Zeek, Elastiflow and NetDisco to provide required management and operations visibility.

​

​

​

​

6 of 20

Research Areas

6

Continuing to use VMware for services (lots of resiliency and features). Bonus was that MSU and UM institutions have covered at least this year’s costs (~$10K); added vSAN capability

WLCG Security Operations Center: Hardware in place, needs Zeek deployment and MISP integration

Using CheckMK for service/host monitoring; now packaged using containers with LE & NGINX

Continue to maintain/upgrade Elasticsearch for syslogging, security and monitoring

Elastiflow deployed monitoring UM in/out flows (utilizes our Elasticsearch)

BOINC: Improves resource use and enables use when grid systems fail or we need to drain for upgrades. Have tweaked / tuned a lot. Could export to other Tier-2s?

Networking: WLCG Site monitoring in place, Flow labeling on dCache hosts, deployed PTP

Considering best options for future provisioning (Cobbler, Foreman, ?); want RHEL9 type OS

Have deployed multi-100G TrueNAS iSCSI storage systems for VMware VMs, replaces unsupported Dell MDxxxx systems at higher performance and lower costs.

​

7 of 20

Run-3 Readiness

7

Item

Site

​

Ready Time

OSG 3.6

Whole Site

OSG 3.6

Apr 8, 2022

HTCondor 9.0.X

Whole Site

9.0.13

Apr 8, 2022

dCache 7.2.x

Whole Site

7.2.16

Apr 21, 2022

IPv6

Whole Site

Dual-stacked

Dec 5, 2019

FY21 Compute Received

UM

10x R6525

Jun 3, 2022

FY21 Compute Received

MSU

8x R6525

Feb 7, 2022

FY22 Compute Received

UM

15x R6525

July 23, 2022

FY22 Storage Received

UM

5x R740xd2

Mar 2, 2022

FY22 Compute Received

MSU

14x R6525

July 23, 2022

FY22 Storage Received

MSU

3x R740xd2

Mar 2, 2022

8 of 20

Backup Slides

Questions or Comments?

8

9 of 20

BOINC work

9

  • BOINC optimization
    • We set up different work node test groups with different configurations for cgroup and BOINC, and through 2 months’ data, the results suggest configuring BOINC as a service and put it under the system.slice cgroup can improve the CPU Efficiency loss for grid jobs by 5%, and on top of that, having BOINC jobs use 50% of the cores (instead of 100%)can further improve the CPU Efficiency loss by another 5%. Following this result, we reconfigured the BOINC services and its cgroup configurations on all the work nodes.
  • Harvest from running BOINC jobs.
    • It increases the CPU Utilization of the cluster, taking the past 5 months for example, the average cores for ATLAS is 13000, and the CPU Utilization reaches 92% combining both Grid and BOINC jobs (75% for Grid, and 17% for BOINC jobs)
    • Fill the cluster during site downtime/cluster draining (HTCondor update)/grid service or network issues as BOINC jobs requires only the work node itself and intermittent network access

10 of 20

Part of Our Problem

10

Cabling: messy, wrong/missing labels, bad airflow, unworkable!

Switch is here! Try to plug in a new cable!!

11 of 20

AGLT2 Plans

11

As noted, AGLT2 has two sites, one at UM, one at MSU. Both are in need of significant upgrade and rework. Our plans:

The MSU site has been located in the Physics building but the infrastructure there is old and fragile (CRACs, space in general)

  • MSU is providing new space in their campus data center to host the AGLT2 equipment and will support the equipment move.
  • Redundant power, networking provided in each rack

The UM site will remain in the LS&A College machine room but needs a complete rewire and reorganization.

All the above work is scheduled for next week (June 14-18, 2021)

12 of 20

Goals for our Network Rewire/Upgrade

12

For AGLT2, our goals is to fix a number of issues

  • Incrementally acquiring servers and network devices created a mess in terms of wiring and airflow
  • Using multiple vendors switches (of different ages) has led to a fragile network that is unable to be optimized, managed and debugged
  • Not have central configuration of our switches has led to mistakes and difficulty in implementing new features
  • Lack of resiliency makes upgrades and config changes difficult, requiring downtimes to make major updates.

13 of 20

New AGLT2 UM LAN

13

The new LAN/WAN design has a border, core and rack-level access on both data and management planes.

Unreached goal: VXLAN, EVPN to each rack (problem between Dell and Cisco :( )

Resiliency from LACP(VLT,MLAG) trunks between redundant switches

14 of 20

UM Switch Testing Rack

14

UM staged all the new LAN/WAN equipment in a testing rack

Allowed us to test and preconfigure switches using Ansible

15 of 20

UM Rewiring Details

15

Our plan is to completely recable our AGLT2 UM site, first removing all existing cabling (power, networking, etc) and then neatly installing new pre-labeled cables

  • RJ45 will use “slim” cables
  • DAC cables will be replaced by fiber optic cables + transceivers
  • Use cable management (horizontal/vertical)
  • Power cables are length optimized
  • We will minimize inter-rack cabling

Have 1 week downtime planned (June 14-18)

​

16 of 20

Ansible and GitHub

16

One of the challenges we have faced since AGLT2 began is maintaining a consistent network configuration.

All maintenance was done via ‘ssh’ to each device. Rancid was used to capture configurations and store in svn

The site refreshes allow us to explore new options.

We have chosen Ansible and Github as tools to support central configuration and revision management.

  • Base switch configurations already exist in Ansible
  • This week we will be loading the planned configs

17 of 20

Network Security

17

AGLT2 has been working with the WLCG SOC effort to help secure our networks while maintaining performance.

Current network has a Zeek+MISP+Elasticsearch setup for dual 40G. Cost to set up was about $2K plus repurposing an R630

​

Our new network is 4x100G

We have purchased two “network capture” nodes (Dell R7525) each with two Mellanox /NVIDIA Bluefield-2 NICs (PCIe Gen4, dual 100G, multiple cores)

Summer students prototyping new systems now

18 of 20

Network Monitoring and Measurement

18

Sites need to make sure they can “see” and understand how their network is performing.

AGLT2 uses a combination of CheckMK, Elasticsearch and custom scripts to monitor our networks.

  • CheckMK provides port traffic/errors/discards (switch/server)
  • Elasticsearch tracks logging from devices and various metrics

For our new network we are deploying multiple 100G perfSONARs, strategically located to cover our border, storage location equivalent and inter-site. All sites should strategically deploy perfSONAR and keep it updated!

19 of 20

Consider PTP in Planning

19

For about $1500, AGLT2 added dual GPS clocks to enable PTP

  • Challenge is the antenna; ideally switch support
  • PTP provides < 1 microsecond time accuracy
  • Makes perfSONAR latency much

more powerful!

20 of 20

Acknowledgements

20

We would like to thank MSU and UM networking personnel, especially Tim DeNike and Nick Grundler for contributing to AGLT2 networking and this presentation.

​

We want to especially acknowledge the National Science Foundation for their support on grant # 162473