1 of 60

Better Living Through Statistics

Monitoring Doesn't (Have To) Suck

Jamie Wilkinson

Site Reliability Engineer, Google

jaq@{spacepants.org,google.com}

@jaqpants

2 of 60

#monitoringsucks

I love monitoring. This hashtag makes me sad.

But what to talk about?

I have no idea what has gone on in the real world since going into the black hole.

These #monitoringsucks people seem to be on the right path... do I have anything new to add?

3 of 60

Validation

http://blog.lusis.org/blog/2012/06/05/monitoring-sucking-just-a-little-bit-less/

"Instead of alerting on data and then storing it as an afterthought (perfdata anyone?) let’s start collecting the

data, storing it and then alerting based on it."

4 of 60

What is monitoring?

Measuring,

Recording,

Alerting,

Visualising

5 of 60

Monitoring systems

automate the boring parts:

Measuring,

Recording,

Alerting,

Visualisation

so you have more time to do fun things... and debug the occasional emergency.

6 of 60

The current state of monitoring

  1. blackbox monitoring resulting in alerts
  2. whitebox monitoring resulting in charts

7 of 60

8 of 60

9 of 60

Blackbox vs Whitebox

Blackbox: treat the system as opaque: you can only use it as a user would.

a.k.a. "probing"

c.f. cucumber-nagios

fuel indicator light

10 of 60

Blackbox vs Whitebox

Whitebox: expose the internal state of the system for inspection

a.k.a. instrumentation, telemetry,

... ROCKET SCIENCE

c.f. ... graphite? new relic? metrics?

tacho, fuel gauge, water meter

11 of 60

What's wrong with blackbox?

Only boolean: no visibility into why

  • Why is the site slow?
  • Why has image serving stopped working?

No predictive capability

  • How long until we need more disks? cpus? datacenters?

All the stuff you do with historical data...

12 of 60

Pseudo-whitebox

Expose some internal state

Test the latest point in time against a threshold.

Return Y/N/maybe?

... ~same as blackbox

e.g. fuel indicator light, again

13 of 60

14 of 60

Problems with "check+alert" model

  • Thresholds vary among instances, tuning difficult.
  • Adding new targets, new checks is lots of effort.
  • Check script logic performs the measurement and the "judgement" all in one.
  • Alerts for things you can't act on.
  • Duplicates
  • Physical resource limits.

15 of 60

So, Why does #monitoringsuck?

TL;DR:

when the cost of maintenance is too high to improve the quality of alerts

16 of 60

An idea...

http://blog.lusis.org/blog/2012/06/05/monitoring-sucking-just-a-little-bit-less/

"Instead of alerting on data and then storing it as an afterthought (perfdata anyone?) let’s start collecting the

data, storing it and then alerting based on it."

17 of 60

18 of 60

19 of 60

20 of 60

Structure of timeseries

21 of 60

Alerting on thresholds

22 of 60

Alert when beer supply low

if cases - 1 - 1 <= 1:

alert BarneyWorriedAboutBeerSupply

23 of 60

Disk full alert

Alert when 90% full

Different filesystems have different sizes

10% of 2TB is 200GB

False positive!

Alert on absolute space, < 500MB

Arbitrary number

Different workloads with different needs: 500MB might not be enough warning

24 of 60

Disk full alert

More general alert:

How long before the disk is full?

and

How long will it take for a human to fix an (almost) full disk?

25 of 60

Alerting on rates of change

26 of 60

Dennis Hopper's Alert

if speed >= 50mph:

alert BombArmed

if BombArmed && speed < 50mph:

...

27 of 60

Keanu's Alert

if speed >= 50mph:

alert SaveTheBus!

28 of 60

29 of 60

Keanu's alert

v - a * t = 50

50 - v = - a * t

(v - 50)/a = t

if (v - 50)/a <= time to save bus:

alert StartSavingTheBus!

30 of 60

New tools at our disposal

Calculus!

the derivative of speed is acceleration

the derivative of acceleration is ... jerk

(impulse?)

31 of 60

Error spike

error count

32 of 60

Rate of errors vs nominal rate

rate of change increases greater than expected

errors per second

33 of 60

calculate rates of change of timeseries

34 of 60

Another new tool

Not just looking at the latest data point, or the derivative at the latest point, otherwise just reimplementing pseudo-whitebox

Look back 5 minutes, 1 hour, 7 days, back to the dawn of time.

Useful minimum: 2.5x the sampling interval.

35 of 60

Traffic spike with threshold

worth getting out of bed for?

36 of 60

Traffic spike with threshold

worth getting out of bed for?

Δt

NO

maybe?

Δt

37 of 60

observe timeseries history to gain context

38 of 60

Timeseries Have Types

Counter: monotonically nondecreasing

"preserves the order" i.e. UP

"nondecreasing" can be flat

39 of 60

Timeseries Have Types

Gauge: everything else... not monotonic

40 of 60

Counters FTW

Δt

41 of 60

Counters FTW

no loss of meaning after sampling

Δt

42 of 60

Gauges FTL

Δt

43 of 60

Gauges FTL

lose spike events shorter than sampling interval

Δt

44 of 60

prefer counters over gauges

45 of 60

Another new tool

Instances in a cluster don't work alone.

Discover the properties of the system by aggregating the parts.

How many queries per second is your cluster receiving?

Take the rate of the counters, and then sum them together!

46 of 60

Aggregation

cluster rate = rate(instance 1 + instance 2)

47 of 60

aggregate to each grouping in the system

48 of 60

high school maths recap

49 of 60

Timeseries Operations: Rates

δ(counter)/δt = gauge

δ(gauge)/δt = gauge

(beware of sampling errors)

50 of 60

Timeseries Operations: Aggregation

Σ0..n(counter) = counter

Σ0..n(gauge) = gauge

(beware of sampling error and quantization)

51 of 60

Timeseries Operations: Ratios

counter / counter = counter: instant means

gauge / gauge = gauge: rate comparisons

e.g. New deployment

δ(errors) / δ(queries) > threshold?

52 of 60

What to measure?

53 of 60

What to measure?

Whatever is important to the business!

Building blocks:

  • Queries per second
    • What is a query?
  • Errors per second
    • Type of error
  • Latency
    • by query type, response code, payload size
  • Bandwidth
    • by direction, query type, response code, ...

54 of 60

What not to measure

  • Load average..

55 of 60

What to alert on?

  • Rate of change of QPS outside normal cycles
  • Ratio of errors to queries
  • Latency (median, 95th percentile) too high
  • Rate of change of bandwidth

Whatever is important to the business!

56 of 60

When to alert?

Alerts are not logs.

Make sure it's ACTIONABLE...

then DOCUMENT IT

3-month-in-the-future-you will thank you.

57 of 60

Blackbox testing still necessary

Blackbox tests are end-to-end tests.

End-to-end testing by definition covers everything you have missed.

You still have charts in the timeseries to inspect, right?

58 of 60

TL;DL

  • Do maths on your timeseries
  • Keep counters instead of gauges, derive rates
  • Compare them to one another
  • Do historical analysis (compare values over time)
  • Alert only when action can be taken

59 of 60

HOW?

Attach a statistical package to your timeseries

database, and experiment

  • R
  • numpy
  • Processing
  • your favourite here

Make smarter alerts!

60 of 60

Questions?