1 of 37

Rolling out

Service Levels

across a large org

Alex Ewerlöf

Senior Staff Engineer

Volvo Cars

alexewerlof.com

2 of 37

What does reliability mean to you?

💭

3 of 37

What does reliability mean to you?

Availability

Scalability

Predictability

Resilience

Performance

Uptime

Security

Trustability

Accessibility

Usability

Reasonability

Traceability

Dependability

Stability

Safety

Consistency

Accuracy

Tolerance

4 of 37

Why do we care?

Loss of opportunities

Seize opportunities

Financial loss

Make revenue

Loss of reputation

Earn trust

Incident panic 😱

Controlled risk

Error budget

5 of 37

Shift in conversation

SLI

Error budget

6 of 37

SLI: Service Level Indicator

SLO: Service Level Objective

SLA: Service Level [legal] Agreement

The metric indicating how your consumers perceive reliability of your service. Percentage of good in a given period.

Where you want your SLI to be. Guides the optimization efforts. For example 99.9%.

Same as SLO but towards paying stakeholders. Breaching SLA leads to punishment, legal action, or compensation.

7 of 37

SLA = SLI + SLO + Agreement

What happens if you don’t??

What to measure?

Minimum Commitment

[legal]

8 of 37

SLI: measures reliability

SLO: sets our objective

9 of 37

SLI

Percentage of good in a given time period

Example:

SLI =

Good

Valid

x 100

SLI =

Number of requests�where latency < 2000ms

Total number of requests over a period

x 100

10 of 37

SLI Example

Let’s imagine two systems. They can be a front-end and backend, or two microservices, or a backend and database.

Depends on responses from

System A

System B

11 of 37

SLI Example

Let’s assume that the analysis of System A shows that System B should return a response in less than 2000ms

The response time should be < 2000ms

System A

System B

12 of 37

Data points

System A

System B

3400 ms

1990 ms

2020 ms

5100 ms

2300 ms

1200 ms

1002 ms

1072 ms

1990 ms

1005 ms

6000 ms

2000 ms

1300 ms

1092 ms

1482 ms

1565 ms

1250 ms

1642 ms

1032 ms

1720 ms

13 of 37

Different readings

Client

More data�More noise�More realistic

Easier to measure

More control�Less realistic

Internet

Edge

Cluster

Instance

3rd party

End user

Cloud provider

1

2

3

4

5

6

14 of 37

Mapped to graph

Threshold of good events

SLO window

15 of 37

SLS: Service level status

SLS =

992300

1000300

x 100 = 99.2%

Number of good RESPONSES

Total number of REQUESTS

Service Level Status

16 of 37

99.2%

Good?

Bad?

17 of 37

Good?

Bad?

Depends on SLO!

SLO

99.5%

99%

SLS

99.2% ❌

99.2% ✅

Right answer✅

18 of 37

Handshake

My Service

My dependencies

🤝

My

Service Consumers

My dependencies

My dependencies

🤝

🤝

🤝

19 of 37

💪

Prepare

My Service

My

Service Consumers

Their

Service Consumers

⏬🤝

20 of 37

My Service

⏬🤝

My

Service Consumers

Their

Service Consumers

⏬🤝

Share the risk

21 of 37

My Service

🔊⏫

My

Service Consumers

Their

Service Consumers

Negotiate

22 of 37

“Get away with”

“An SLO should define the lowest level of reliability that you can get away with for each service.”

— Jay Judkowitz and Mark Carter (Google PMs)

23 of 37

Example:

  • 99% is better than 99.4%
  • 97% is better than 99%

IF YOU CAN GET AWAY WITH THAT!

Underpromise and overdeliver

(not the other way around!)

24 of 37

Does anyone use SLOs?

25 of 37

Error budget

…complements SLO

Example:

SLO = 99%

Error Budget = 100% - 99% = 1%

26 of 37

Service Consumer

System Architecture

Service Provider

Owns

Consumes

Owned by service provider

Service

System/Component

27 of 37

Service Consumer

System Architecture

Service Provider

Owns

Consumes

Owned by service provider

Service

System/Component

No one’s business!

28 of 37

SLI workshop

Intro(45-60m)

Workshop(90-180m)

What is SLI/SLO/SLA and what’s expected of teams?

Risks(10-20m)

Metrics(20-30m)

SLIs(10-20m)

SLOs(10-20m)

(10m break)

Building the mental model and tools to set service levels and own them.

Find meaningful SLIs Commit to reasonable SLOs

Define Service(10-20m)

29 of 37

Adoption steps

Step 1

Intro & workshop

Step 2

Written service level document

Step 3

Measure reliability in real time

Step 4

Alerting & On-call

30 of 37

Service Level Calculator

31 of 37

User facing surface

User

Systems and their dependencies

32 of 37

User

System ownership

Org chart

33 of 37

So we need to map a graph to tree

Org chart

Architecture Diagram

Systems, components and their dependencies

Teams, people and their reporting lines

34 of 37

User

SLI/SLO

SLI/SLO

SLI/SLO

SLI/SLO

SLI/SLO

SLI/SLO

SLI/SLO

35 of 37

SPOG

Incidents

Service Level Status

Org chart

SLS: 95.1% Incidents: 4

SLS: 100% Incidents: 0

SLS: 97.2% Incidents: 3

SLS: 95% Incidents: 1

SLS: 95% Incidents: 0

SLS: 90.2% Incidents: 2

SLS: 93.2% Incidents: 0

SLS: 94.7% Incidents: 1

SLS: 80.8% Incidents: 1

SLS: 98% Incidents: 0

SLS: 99% Incidents: 0

SLS: 93% Incidents: 0

SEV 1

SEV 2

SEV 3

SEV 4

SLO

Last 30 days

36 of 37

Recap

  • Service levels set expectations
  • SLI measures reliability from consumer’s perspective
  • Adopt SLOs in steps
  • Error budgets balance innovation vs risk
  • Reliability should map to accountability

37 of 37

Online book

on Substack