1 of 19

From Fastnet Lighthouse to the Cloud:Dovetailing Resilience into Greener Data Centre Infrastructure

James Delaney, MSc Cloud Computing (MTU), Data Centre Engineer (UCC)

Supervisor: Dr Patrick McCarthy (MTU)

2 of 19

Introduction

A software-defined approach inspired by the dovetail toggle of the Fastnet Lighthouse

Making Kubernetes clusters more resilient and energy-efficient, and stewarding a greener data-centre future

3 of 19

Background – The Fastnet Lighthouse Analogy

  • Why the Fastnet?

  • What makes its design uniquely resilient, and relevant to data centres?

  • Preserving fault tolerance where Kubernetes falters – resilience through interlocking design

Background –

The Fastnet Analogy

4 of 19

Background – The Fastnet Lighthouse Analogy

  • Why the Fastnet?

  • What makes its design uniquely resilient, and relevant to data centres?

  • Preserving fault tolerance where Kubernetes falters – resilience through interlocking design

Background –

The Fastnet Analogy

5 of 19

Solving the Data Centre Resilience Problem for Sustainability

  • Data centres are the main driver of Ireland’s electricity growth since 2015

  • Demand from data centres grew by 22.6% per year (vs 0.4% in other sectors)

  • Ireland’s electricity use rose 24.7% (2012–2022) – 2nd fastest in the EU, while EU demand fell 3.1%

  • Household electricity use stayed flat → growth driven by data centres

  • Similar trend worldwide, with data centres among the fastest-growing energy consumers

H. Daly, “Data centres in the context of Ireland’s carbon budgets,” 2024.

6 of 19

Telemetry

How can real-time power telemetry from Kepler be integrated into RL-driven scheduling for energy-aware workload placement?

Resilience

How can a RL-based workflow orchestration framework, leveraging the dovetail toggle approach, balance energy efficiency, scalability, and fault tolerance in large-scale data centres?

Efficiency

How can Reinforcement Learning (RL) be applied to optimise energy efficiency in Kubernetes orchestration?

Research Questions

7 of 19

System Architecture Design

  • Defined six-layer modular framework (data → analysis → recommendation → RL → migration → reporting).
  • Designed dovetail scheduling logic to preserve ReplicaSet fault tolerance.

Implementation & Environment

  • Developed in Python (v3.10) and deployed on k3s/k3d simulated clusters.
  • Integrated Docker, Kepler, and Prometheus for real-time telemetry.
  • PostgreSQL database for persistent metric logging.

Reinforcement Learning Agent

  • Tabular Q-learning with 10,000 training episodes.
  • State space: CPU/memory levels + container mappings.
  • Reward: balance energy efficiency, consolidation, and fault tolerance.

Analysis & Validation

  • Visual telemetry in Grafana dashboards.
  • Human-readable logs and rationales for transparency.

Evaluation Procedure

  • Compared RL against advanced rule-based, random, and default Kubernetes schedulers.
  • Tested across 4, 8, and 12 node clusters.
  • Key metrics: energy, CO₂ savings, migrations.

Methodology

8 of 19

Research Design –

The Dovetail Toggle

  • Designed by William Douglass

  • 2074 granite stones – each one hand carved

  • Interlocking blocks joined both vertically and horizontally

  • Each stone knitted in with the one above, below, and either side

  • No single stone could fail in isolation – to crack it would mean bringing down the entire edifice

9 of 19

Research Design –

The Dovetail Toggle

  • Resilience through interlocking structural design

  • Translated into a distributed computing context

  • Each node depends on, and protects its neighbours

10 of 19

Active Nodes

Cold Standby (Idle) Nodes

Primary containers (hosted on this node)

Replica of primary containers from previous node

Research Design –

The Dovetail Toggle

  • Nodes interlock laterally via primaries and replicas

  • Circular failover ensures no replica shares its primary node

  • Cold standbys activate only when required for scaling or node failure

  • Engineered for resilience, optimised for sustainability

11 of 19

Active Nodes

Cold Standby (Idle) Nodes

Primary containers (hosted on this node)

Replica of primary containers from previous node

Research Design –

The Dovetail Toggle

Dovetail Migration Logic Summary

  • Each node hosts its own primaries and replicas of the previous node’s primaries

  • When a node shuts down:
    • Primaries → move two nodes forward
    • Replicas → move one node forward

  • Maintains circular fault tolerance and avoids co-location of primaries and replicas

A + F

B + F

12 of 19

System Level Architecture

6

Reporting Systems

Delivers real-time and persistent reporting outputs, including detailed operator-facing rationales.

1

Data Collection Layer

Gathers real-time CPU, memory and power metrics from containers and nodes via Docker, Kepler, and Prometheus.

2

Analysis Layer

Processes collected data to assess resource utilisation, energy consumption, and operational anomalies.

3

Recommendation Engine

Generates actionable migration suggestions based on multi-criteria optimisation.

5

Migration Planner

Designs migration sequences that maintain ReplicaSet fault tolerance, based on the dovetail toggle architecture.

4

RL Module

Trains a Q-learning agent to improve container placement strategies dynamically.

13 of 19

Implementation & Measurements

  • Energy models calibrated to Dell PowerEdge R740 specs

  • Idle draw assumed at 36% of peak (Berkeley Lab 2024)

  • Energy & CO₂ estimated via SEAI emission factor (0.2548 kg CO₂ / kWh)

  • Evaluation under identical 4, 8, and 12 node clusters

  • Live telemetry confirmed post-migration energy drops in Grafana

Implementation & Measurements

14 of 19

Implementation & Measurements

  • Energy models calibrated to Dell PowerEdge R740 specs

  • Idle draw assumed at 36% of peak (Berkeley Lab 2024)

  • Energy & CO₂ estimated via SEAI emission factor (0.2548 kg CO₂ / kWh)

  • Evaluation under identical 4, 8, and 12 node clusters

  • Live telemetry confirmed post-migration energy drops in Grafana

Implementation & Measurements

Progressive Reduction in Daily CO₂ Emissions During RL Agent Training Across Cluster Sizes

Real-time Grafana dashboard showing reduced power usage after three nodes were powered down in a 12-node cluster.

15 of 19

Reinforcement Learning Policy

(this research)

Nodes Shutdown: 3

Annual Savings (kWh): 2522.88

co2 Reduction (kg): 642.84

Advanced Heuristics Policy

Nodes Shutdown: 2

Annual Savings (kWh): 1681.92

co2 Reduction (kg): 428.56

K8s Default &

Random Policies

Nodes Shutdown: 0

Annual Savings (kWh): 0

co2 Reduction (kg): 0

Rule Based Policy

Nodes Shutdown: 0.003

Annual Savings (kWh): 2.52

co2 Reduction (kg): 0.64

Findings

12 Node Cluster Simulation

SEAI factor: 0.2548 kg CO2/kWh

Idle power = 36% of peak (Berkeley Lab, 2024)

16 of 19

Reinforcement Learning Policy

(this research)

Nodes Shutdown: 3

Annual Savings (kWh): 2522.88

co2 Reduction (kg): 642.84

Advanced Heuristics Policy

Nodes Shutdown: 2

Annual Savings (kWh): 1681.92

co2 Reduction (kg): 428.56

K8s Default &

Random Policies

Nodes Shutdown: 0

Annual Savings (kWh): 0

co2 Reduction (kg): 0

Rule Based Policy

Nodes Shutdown: 0.003

Annual Savings (kWh): 2.52

co2 Reduction (kg): 0.64

Findings

12 Node Cluster Simulation

SEAI factor: 0.2548 kg CO2/kWh

Idle power = 36% of peak (Berkeley Lab, 2024)

17 of 19

Why This Matters

2 return flights from Cork to London per person

2 months of electricity for an average Irish household

What 25 trees absorb in an entire year

18 of 19

Conclusions & Future Research

Conclusions

  • RL transformed Kubernetes into a self-optimising, energy-aware platform
  • Dovetail design delivered resilience and sustainability in one architecture
  • Outperformed Kubernetes ReplicaSets – kept structure, cut power
  • Lightweight Q-learning achieved real-time, scalable efficiency
  • Marks a progressive step toward carbon-intelligent cloud infrastructure

Future Work

  • Validate performance at scale in a live Kubernetes cluster to measure real-world energy and CO₂ savings
  • Stress-test resilience under live failure conditions to refine self-healing behaviour
  • Adapt the RL architecture for broader orchestration contexts – from edge to virtualised and HPC environments
  • Take the next step toward embedding carbon awareness into native orchestration
  • Enable others to reduce their CO₂ footprint through shared, open innovation

19 of 19

Thank

You