1 of 11

Your drives are

dying right now.

You just don't know which ones.

Aroa Xinping

Data Scientist & ML Engineer

Predictive Analytics for Infrastructure

2 of 11

Agenda

01

The problem

What a single drive failure really costs

02

The opportunity

Your drives already broadcast distress signals

03

Solution

What we built and what it detects

04

The savings

$49,200 per quarter, the math

05

Live output

Real predictions on real drives

06

Under the hood

Why you can trust the model

07

Implementation

One cron job. Zero new hardware

1

3 of 11

THE PROBLEM

A drive costs $30.

Missing its failure costs $10,000.

Drive fails

→

RAID degrades

→

Rebuild stresses

remaining drives

→

Second failure:

TOTAL DATA LOSS

GitLab, 2017: cascading RAID failure. 300 GB of production data lost. 18 hours of downtime.

​

SLA breach penalties, compliance fines, and customer churn: a single failure can cost 6 figures.

Yet: the industry standard is still reactive: replace the drive after it fails.

2

"Industry estimates (Gartner, ITIC). Actual costs vary by infrastructure."

​

​

4 of 11

THE OPPORTUNITY

Your drives already broadcast

distress signals. For free.

SMART sensors — every modern drive monitors its own health: reallocated sectors, uncorrectable errors, temperature, power-on hours. This data is already being collected on your servers. It's just not being used.

318K

drives analysed

28M

daily readings

90

days of data

Source: Backblaze — public production data from real data centers

3

5 of 11

SOLUTION

Build a model that detects

96% of drive failures before they happen.

96%

detection rate

(205 of 213 failures)

9

false alarms

($450 total cost)

0.3%

failure rate in data

(extreme imbalance)

4

6 of 11

For the most technical audience:

Left: 99% of failures with almost zero false alarms

​

Right: even with only 0.3% failure rate, precision stays above 95%

"The gray line guesses. The green line knows."

5

7 of 11

$49,200 saved

per quarter.

Don't optimise for statistics. We optimise for your bottom line.

Default model

$80,450

total cost / 205 detected / 9 false alarms

Cost-optimised

$31,250

total cost / 210 detected / 25 false alarms

16 extra false alarms × $50 = $800

5 extra failures detected × $10,000 = $50,000 saved

6

8 of 11

LIVE OUTPUT

$ predict.py

DRIVE

PROB

RISK

ACTUAL

ZA17XZ8Z

100.00%

HIGH

FAILED

ZA12YYNF

100.00%

HIGH

FAILED

ZA171BTC

0.00%

LOW

healthy

ZA18CD79

0.00%

LOW

healthy

74W0A0T4

0.00%

LOW

healthy

5/5 correct

threshold: 0.04 (cost-optimised)

One command.

Real drives.

Real predictions.

This is not a simulation.

These are actual production drives

from Backblaze data centers.

github.com/aroaxinping/

drive-failure-predictor

try it yourself!

7

9 of 11

UNDER THE HOOD

Why you can trust this.

Trained on real data. 318K production drives, not synthetic benchmarks

Validated across time. Train on Jan-Feb, test on March. Not a random split that leaks information.

Optimised for your costs. Tuned for business impact ($10K per miss vs $50 per false alarm), not academic metrics

Interpretable. The model learns what your engineers already know: degradation speed matters more than current error count.

8

10 of 11

IMPLEMENTATION

One cron job. Zero new hardware.

smartmontools

reads SMART sensors daily (already on most Linux servers)

→

Model scores

each drive gets a

failure probability

→

Alert

high-risk drives flagged

in Grafana / PagerDuty

→

Replace

scheduled replacement

in the next maintenance window

What this doesn't require

  • No new hardware or sensors
  • No changes to existing infrastructure
  • No dedicated team:

one engineer, one afternoon

Next steps

  • Multi-quarter training for seasonal patterns
  • Per-model failure profiles (Seagate vs WDC)
  • Remaining useful life estimation

9

11 of 11

Every quarter you wait,

another $50K walks out the door.

Let's talk about your infrastructure.

Aroa Xinping

aroaxinping@gmail.com

github.com/aroaxinping/drive-failure-predictor