Your drives are
dying right now.
You just don't know which ones.
Aroa Xinping
Data Scientist & ML Engineer
Predictive Analytics for Infrastructure
Agenda
01
The problem
What a single drive failure really costs
02
The opportunity
Your drives already broadcast distress signals
03
Solution
What we built and what it detects
04
The savings
$49,200 per quarter, the math
05
Live output
Real predictions on real drives
06
Under the hood
Why you can trust the model
07
Implementation
One cron job. Zero new hardware
1
THE PROBLEM
A drive costs $30.
Missing its failure costs $10,000.
Drive fails
→
RAID degrades
→
Rebuild stresses
remaining drives
→
Second failure:
TOTAL DATA LOSS
GitLab, 2017: cascading RAID failure. 300 GB of production data lost. 18 hours of downtime.
SLA breach penalties, compliance fines, and customer churn: a single failure can cost 6 figures.
Yet: the industry standard is still reactive: replace the drive after it fails.
2
"Industry estimates (Gartner, ITIC). Actual costs vary by infrastructure."
THE OPPORTUNITY
Your drives already broadcast
distress signals. For free.
SMART sensors — every modern drive monitors its own health: reallocated sectors, uncorrectable errors, temperature, power-on hours. This data is already being collected on your servers. It's just not being used.
318K
drives analysed
28M
daily readings
90
days of data
Source: Backblaze — public production data from real data centers
3
SOLUTION
Build a model that detects
96% of drive failures before they happen.
96%
detection rate
(205 of 213 failures)
9
false alarms
($450 total cost)
0.3%
failure rate in data
(extreme imbalance)
4
For the most technical audience:
Left: 99% of failures with almost zero false alarms
Right: even with only 0.3% failure rate, precision stays above 95%
"The gray line guesses. The green line knows."
5
$49,200 saved
per quarter.
Don't optimise for statistics. We optimise for your bottom line.
Default model
$80,450
total cost / 205 detected / 9 false alarms
Cost-optimised
$31,250
total cost / 210 detected / 25 false alarms
16 extra false alarms × $50 = $800
5 extra failures detected × $10,000 = $50,000 saved
6
LIVE OUTPUT
$ predict.py
DRIVE
PROB
RISK
ACTUAL
ZA17XZ8Z
100.00%
HIGH
FAILED
ZA12YYNF
100.00%
HIGH
FAILED
ZA171BTC
0.00%
LOW
healthy
ZA18CD79
0.00%
LOW
healthy
74W0A0T4
0.00%
LOW
healthy
5/5 correct
threshold: 0.04 (cost-optimised)
One command.
Real drives.
Real predictions.
This is not a simulation.
These are actual production drives
from Backblaze data centers.
github.com/aroaxinping/
drive-failure-predictor
try it yourself!
7
UNDER THE HOOD
Why you can trust this.
Trained on real data. 318K production drives, not synthetic benchmarks
Validated across time. Train on Jan-Feb, test on March. Not a random split that leaks information.
Optimised for your costs. Tuned for business impact ($10K per miss vs $50 per false alarm), not academic metrics
Interpretable. The model learns what your engineers already know: degradation speed matters more than current error count.
8
IMPLEMENTATION
One cron job. Zero new hardware.
smartmontools
reads SMART sensors daily (already on most Linux servers)
→
Model scores
each drive gets a
failure probability
→
Alert
high-risk drives flagged
in Grafana / PagerDuty
→
Replace
scheduled replacement
in the next maintenance window
What this doesn't require
one engineer, one afternoon
Next steps
9
Every quarter you wait,
another $50K walks out the door.
Let's talk about your infrastructure.
Aroa Xinping
aroaxinping@gmail.com
github.com/aroaxinping/drive-failure-predictor