1 of 24

Data Center Life Cycle Management @ Scale

Life cycle management for hyperscale introduces unique challenges for manageability. This presentation showcases a blueprint for a data model which can be aggregated across various levels and stages of hyperscale evolution.

2 of 24

Data Center Life Cycle Management @ Scale

Nirav Shah, Cloud Software Architect, Intel

Jim Harford, System Architect, Broadcom

Scott Ramsey, Technologist, Dell Technologies

SUSTAINABLE SCALABLE COMPUTATIONAL INFRASTRUCTURE

HW MGMT

3 of 24

CSM WG Overview

Purpose

  • Define comprehensive at scale remote service model (inclusive of different sizes)
  • Standardize interfaces for at scale remote service model & enable added services
  • Deliver services across the boundaries of ownership
  • Ability to integrate vendor tools using common framework
  • Influence spans across multiple OCP disciplines

4 of 24

Life Cycle Opportunities 🡪 Recap

Plan/Design

1

Procure/Deploy

2

Operate

3

Decommission

4

  • Real Estate
  • Utilities
  • Network
  • Systems sourcing
  • Configuration
  • Conformance/Certification
  • Validation
  • Orchestration/Ramp
  • Utilization
  • Fault Detection/Repair
  • Servicing
  • Inventory management
  • Backup & Reset
  • Decommission
  • Recycle

STAGES

FACTORS

STANDARDS

5 of 24

Data Center @ scale

10s of thousands of geo located and/or distributed servers!

Racks!

Racks!

Rack 1

Rack N (geo-located)

Rack Z(distributed)

  • A universal data model across components, systems and physical/virtual aggregates.

Physical Aggregate

Virtual Aggregate

6 of 24

Various Phases

  • Interconnect the phases
  • Interconnect the use cases.

7 of 24

Utilization

Utilization Use cases

  • Define power budget & optimal utilization / power ratio.
  • Identify system/aggregate configurations meeting power and utilization budgets
  • Identify components vendors supporting configurations and budgets.
  • Measure utilization for Debuggability, efficiency and/or metering
  • Measure average performance o/ power consumption for TCO
  • Measure conformance to spec’d performance targets.
  • Detect thermal margining overload
  • Detect power margining overloads.
  • Compare the target utilization to operation

Plan/Design

1

Procure/Deploy

2

Operate

3

Decommission

4

STAGES

8 of 24

Utilization Requirements

  • Utilization 🡪 snapshot of component/system/aggregate utilization.
  • Power utilization 🡪 snapshot of power utilized by component/system/aggregate
  • Thermal utilization 🡪 snapshot of operating temperature of component/system/aggregate
  • Uptime 🡪 time in seconds since the last reset/power cycle
  • Average utilization 🡪 Utilization averaged over uptime (set to 0 on reset/power cycle)
  • Average power utilization 🡪 power utilization averaged over uptime. (set to 0 on reset/power cycle)
  • Average thermal utilization 🡪 thermal output averaged over uptime. (set to 0 on reset/power cycle)
  • Average Uptime 🡪 Uptime in seconds averaged over the # of resets
  • Power threshold 🡪 Highest power utilization above which component/system/aggregate “may” be throttled/reset.
  • Thermal threshold 🡪 Highest temperate above which component/system/aggregate “may” be power-cycled.
  • Utilization threshold 🡪 Lowest utilization at which component/system/aggregate “may” not-operate or sleep.

9 of 24

Health Status

Health Status Use cases

  • Define availability/redundancy targets
  • Define fault tolerance targets
  • Define fault recover/sparing targets
  • Identify suppliers/configurations that meet availability, fault tolerance and recovery targets
  • Measure correctable and uncorrectable faults against thresholds
  • Track health status to prevent failures.
  • Spare failed or about to fail hardware
  • Identify patterns of failures, proactively address those failures.
  • Backup and offline faulty hardware
  • Postmortem analysis
  • Sustainably decommission fault hardware.

Plan/Design

1

Procure/Deploy

2

Operate

3

Decommission

4

10 of 24

Health Requirements

Fault/Error

  • Failure Id 🡪 256 bit Unique ID allocated to any new failure with distinct symptoms
  • Type 🡪 Correctable or uncorrectable but not fatal or fatal
  • Count 🡪 Number of error occurrences since component/system/aggregate installed
  • Fix Status 🡪 Unknown/Fixed/Not available/Available/Corrected
  • Telemetry Signature 🡪 Log of a unique set of events leading up to the error.

Health

  • Correctable error # 🡪 Count of all corrected errors since component/system/aggregate installed.
  • Fatal error # 🡪 Count of all fatal errors since component/system/aggregate installed
  • Uncorrectable non fatal error # 🡪 Count of all uncorrectable non fatal errors since component/system/aggregate installed.
  • Health Score 🡪 A positive decreasing # representing the current health of a component/system/aggregate
  • Health Warning Threshold 🡪 Threshold below which component/system/aggregate operates sub optimally
  • Health Critical Threshold 🡪 Threshold below which component/system/aggregate “may” break down

​

11 of 24

Configuration

Network Configuration Use Cases {Example}

  • Define required NIC port speeds and partitioning of PCIe PFs among NIC ports
  • Define TX traffic shaping characteristics for multiple classes of service to be used on NICs
  • Upgrade NIC firmware, driver(s), and/or SW tool(s)
  • Configure network & link parameters needed for basic connectivity to the network
  • Configure TX traffic shaping characteristics for multiple classes of service to be used on NIC
  • Optimize performance via adjustments to BIOS settings, NIC resource allocation, flow steering, and pinning of interrupts to specific CPUs

Plan/Design

1

Procure/Deploy

2

Operate

3

Decommission

4

12 of 24

Configuration Requirement {Example}

  • Initiation & monitoring of upgrades to NIC firmware, driver(s), and SW tool(s)
  • Configuration of NIC hardware parameters
  • Configuration of server’s BIOS settings
  • Download of scripts to local storage used by OS running on server
  • Execution of configuration scripts by OS running on server
  • Discover results from other (opaque to OCP) configuration mechanisms
    • For example: switch->NIC neighbor configuration via LLDP

​

13 of 24

Performance

Performance Use cases

  • Determine benchmarks for off-line performance testing of components/system/aggregates.
    • “micro-benchmarks” (e.g. netperf, all-gather, fio, SPEChpc)
    • synthetic workloads that resemble the demands of production workloads
  • Run off-line performance benchmarks on properly configured elements
  • Measure and record performance statistics while running production workloads
  • Compare performance results of off-line benchmarks versus production workloads

Plan/Design

1

Procure/Deploy

2

Operate

3

Decommission

4

14 of 24

Performance Requirement

  • Availability of performance “micro-benchmarks”
  • Availability of synthetic workloads relevant to a given data center or cloud environment
  • Availability of aggregate level benchmarks to simulate a data center environment.
  • Installation, initiation, and monitoring of benchmark programs on server(s)
  • Configuration of any HW parameters that affect performance
  • Download of test scripts to local storage used by OS running on server
  • Execution of test scripts by OS running on server
  • Access to log files or other artifacts containing measured performance statistics
  • Integrate the logs to a system level telemetry for a customizable view.
  • Configurability at component/system/aggregate level for the scale of telemetry.

15 of 24

Summary

  • Data centers growing at an exponential scale
  • Data center already at a hyperscale introduces unique challenges.
  • Not just the scale but the required sophistication presenting a need to model the life cycle.
  • Identify opportunities for additional needed standardization
  • Walk through scenarios to present unique use cases and requirements.
  • Invite industry to help define the framework that meets these requirements.

​

16 of 24

Call to Action

​

17 of 24

Call to Action

​

18 of 24

Thank you!

19 of 24

Please use one of these membership logos to designate your company’s membership level.

20 of 24

Please use one of these logos if you or your supplier is an OCP Solution Provider.

21 of 24

Please use one of these logos if your organization is hosting an �OCP recognized Innovation Village.

Please use this logo if your Facility is an OCP Ready™ facility.

​

Please use one of these logos if your product has been recognized as an �OCP certified product.

Please use this logo if your product has been approved through the �OCP S.A.F.E. program.

22 of 24

Please use the official track name where applicable.

ARTIFICIAL INTELLIGENCE (AI)

EMEA DEPLOYMENTS

DC SUSTAINABILITY

POWER AND COOLING

SECURITY AND DATA PROTECTION

SPECIAL FOCUS: OPTICS

SPECIAL FOCUS: OPENRAN

SPECIAL FOCUS: QUANTUM

SUSTAINABLE SCALABLE COMPUTATIONAL INFRASTRUCTURE

23 of 24

Please use the appropriate icon representing the Regional Project Group.

24 of 24

Please use the appropriate icon representing the Project Group.

DC FACILITIES

HW MGMT

SECURITY

OPEN SYS FW

TELCO

AI

OPTICS

OPEN CHIPLET

DEPLOYMENTS

NETWORKING

STORAGE

RACK & POWER

SERVER

COOLING ENV

SUSTAINABILITY

TIME APPLIANCES

CMS