Rolling out
Service Levels
across a large org
Alex Ewerlöf
Senior Staff Engineer
Volvo Cars
alexewerlof.com
What does reliability mean to you?
💭
What does reliability mean to you?
Availability
Scalability
Predictability
Resilience
Performance
Uptime
Security
Trustability
Accessibility
Usability
Reasonability
Traceability
Dependability
Stability
Safety
Consistency
Accuracy
Tolerance
Why do we care?
Loss of opportunities
Seize opportunities
Financial loss
Make revenue
Loss of reputation
Earn trust
Incident panic 😱
Controlled risk
Error budget
Shift in conversation
SLI
Error budget
SLI: Service Level Indicator
SLO: Service Level Objective
SLA: Service Level [legal] Agreement
The metric indicating how your consumers perceive reliability of your service. Percentage of good in a given period.
Where you want your SLI to be. Guides the optimization efforts. For example 99.9%.
Same as SLO but towards paying stakeholders. Breaching SLA leads to punishment, legal action, or compensation.
SLA = SLI + SLO + Agreement
What happens if you don’t??
What to measure?
Minimum Commitment
[legal]
SLI: measures reliability
SLO: sets our objective
SLI
Percentage of good in a given time period
Example:
SLI =
Good
Valid
x 100
SLI =
Number of requests�where latency < 2000ms
Total number of requests over a period
x 100
SLI Example
Let’s imagine two systems. They can be a front-end and backend, or two microservices, or a backend and database.
Depends on responses from
System A
System B
SLI Example
Let’s assume that the analysis of System A shows that System B should return a response in less than 2000ms
The response time should be < 2000ms
System A
System B
Data points
System A
System B
…
3400 ms
1990 ms
2020 ms
5100 ms
2300 ms
1200 ms
1002 ms
1072 ms
1990 ms
1005 ms
6000 ms
2000 ms
1300 ms
1092 ms
1482 ms
1565 ms
1250 ms
1642 ms
1032 ms
1720 ms
…
Different readings
Client
More data�More noise�More realistic
Easier to measure
More control�Less realistic
Internet
Edge
Cluster
Instance
3rd party
End user
Cloud provider
1
2
3
4
5
6
Mapped to graph
Threshold of good events
SLO window
SLS: Service level status
SLS =
992300
1000300
x 100 = 99.2%
Number of good RESPONSES
Total number of REQUESTS
Service Level Status
99.2%
Good?
Bad?
Good?
Bad?
Depends on SLO!
SLO
99.5%
99%
SLS
99.2% ❌
99.2% ✅
Right answer✅
Handshake
My Service
My dependencies
🤝
My
Service Consumers
My dependencies
My dependencies
🤝
🤝
🤝
💪
Prepare
My Service
My
Service Consumers
Their
Service Consumers
⏬🤝
My Service
⏬🤝
My
Service Consumers
Their
Service Consumers
⏬🤝
Share the risk
My Service
🔊⏫
My
Service Consumers
Their
Service Consumers
Negotiate
“Get away with”
“An SLO should define the lowest level of reliability that you can get away with for each service.”
— Jay Judkowitz and Mark Carter (Google PMs)
Example:
IF YOU CAN GET AWAY WITH THAT!
Underpromise and overdeliver
(not the other way around!)
Does anyone use SLOs?
✋
Error budget
…complements SLO
Example:
SLO = 99%
Error Budget = 100% - 99% = 1%
Service Consumer
System Architecture
Service Provider
Owns
Consumes
Owned by service provider
Service
System/Component
Service Consumer
System Architecture
Service Provider
Owns
Consumes
Owned by service provider
Service
System/Component
No one’s business!
SLI workshop
Intro�(45-60m)
Workshop�(90-180m)
What is SLI/SLO/SLA and what’s expected of teams?
Risks�(10-20m)
Metrics�(20-30m)
SLIs�(10-20m)
SLOs�(10-20m)
(10m break)
Building the mental model and tools to set service levels and own them.
Find meaningful SLIs Commit to reasonable SLOs
Define Service�(10-20m)
Adoption steps
Step 1
Intro & workshop
Step 2
Written service level document
Step 3
Measure reliability in real time
Step 4
Alerting & On-call
Service Level Calculator
User facing surface
User
Systems and their dependencies
User
System ownership
Org chart
So we need to map a graph to tree
Org chart
Architecture Diagram
Systems, components and their dependencies
Teams, people and their reporting lines
User
SLI/SLO
SLI/SLO
SLI/SLO
SLI/SLO
SLI/SLO
SLI/SLO
SLI/SLO
SPOG
Incidents
Service Level Status
Org chart
SLS: 95.1% Incidents: 4
SLS: 100% Incidents: 0
SLS: 97.2% Incidents: 3
SLS: 95% Incidents: 1
SLS: 95% Incidents: 0
SLS: 90.2% Incidents: 2
SLS: 93.2% Incidents: 0
SLS: 94.7% Incidents: 1
SLS: 80.8% Incidents: 1
SLS: 98% Incidents: 0
SLS: 99% Incidents: 0
SLS: 93% Incidents: 0
SEV 1
SEV 2
SEV 3
SEV 4
SLO
Last 30 days
Recap
Online book
on Substack