1 of 47

Application Monitoring with �Micrometer &�Prometheus &

Grafana

Andrew Fitzgerald�Sonatype

@fitzoh

https://github.com/fitzoh/micrometer-prometheus-grafana-talk

2 of 47

About me

  • JVM + Devopsish
  • Senior Engineer @ Sonatype (Hiring! Lots of remote positions)
  • Rugby
  • Baby!

3 of 47

Agenda

  • 3 Pillars of Observability
  • Micrometer: generate the metrics
  • Prometheus: store the metrics
  • Grafana: view the metrics
  • Demo

4 of 47

3 Pillars of Observability�Don’t know who to credit, sorry ¯\_(ツ)_/¯

5 of 47

Logging

6 of 47

(Structured) Logging

7 of 47

(Structured) Logging

8 of 47

(Structured) Logging

  • If you’re grepping logs, don’t bother
  • If you’re using a log aggregator, do it
  • Logstash logback encoder
  • Example/links in repo

9 of 47

Metrics

  • RED/USE
  • Latency
  • Throughput
  • Error rates
  • Current state
    • Active users
    • Open db connections
    • CPU
    • Disk space

10 of 47

Tracing

  • End to end view of a single request

https://zipkin.io/

11 of 47

Tracing

  • Instrumentation
    • Spring Cloud Sleuth
    • Zipkin
    • OpenTracing
  • Collector
    • Zipkin
    • Jaeger
    • Haystack
    • Commercial offerings (DataDog, New Relic, GCP StackDriver Trace, Elastic, Honeycomb)

12 of 47

Micrometer

13 of 47

Think SLF4J, but for metrics

14 of 47

One API, many backends

  • AppOptics
  • Azure Monitor
  • Atlas
  • AWS CloudWatch
  • Datadog
  • Dynatrace
  • Elastic
  • Ganglia
  • GCP StackDriver
  • Graphite
  • Humio
  • Influx/Telegraf
  • JMX
  • KairosDB
  • New Relic
  • Prometheus
  • SignalFx
  • StatsD
  • Wavefront

15 of 47

Out of the box instrumentation

16 of 47

Out of the box instrumentation

17 of 47

Dimensional metrics

Hierarchical:

  • http_get_requests_count
  • http_200_response_count
  • http_200_path_/user/{id})count

Dimensional:

  • http.server.requests .count (status=200, method=GET, path=/user/{id})

18 of 47

Meters

A Meter is uniquely defined by its name and dimensions/tags

  • Name: dot delimited string
    • http.server.requests
    • hikaricp.connections
  • Dimensions/Tags: set of key/value pairs
    • method=get
    • status=OK
    • path=/user/{id}
  • Description (optional)
  • Unit (optional, not applicable for timers)

19 of 47

Registries

  • Registries create and store Meters
  • Each metric backend has a Registry
    • Add a meter to a Prometheus registry to publish that metric to Prometheus
  • Composite registries (and the global Registry) can be used to publish to multiple registries/backends
  • Can apply common tags at the registry level (environment, version, region)
  • Registries can programmatically filter/modify included Meters

20 of 47

Metric types

  • Counters
    • Number of requests

21 of 47

Counter

22 of 47

Metric types

  • Counters
    • Number of requests
  • Gauges
    • Current open DB connections

23 of 47

Gauge

24 of 47

Metric types

  • Counters
    • Number of requests
  • Gauges
    • Current open DB connections
  • Timers
    • HTTP request timer

25 of 47

Timer

26 of 47

Metric types

  • Counters
    • Number of requests
  • Gauges
    • Current open DB connections
  • Timers
    • HTTP request timer
  • Long Task Timers
    • Long running batch job

27 of 47

Long Task Timer

28 of 47

Metric types

  • Counters
    • Number of requests
  • Gauges
    • Current open DB connections
  • Timers
    • HTTP request timer
  • Long Task Timers
    • Long running batch job
  • Distribution Summaries
    • Request payload sizes

29 of 47

Distribution Summary

30 of 47

Distribution Statistics

Timers and Distribution Summaries can optionally emit distribution statistics

  • Percentiles: Micrometer will precompute specified percentile statistics
    • “How long does it take for 95% of requests to finish?”
  • SLA:
    • “How many requests took longer than 100ms?”
  • Percentile Histogram
    • “Give me a full histogram with 276 buckets of timing information and I’ll figure out what to do with it later”

31 of 47

Distribution Statistics

32 of 47

Prometheus

33 of 47

Prometheus

https://prometheus.io/docs/introduction/overview/

34 of 47

Pull model

35 of 47

Exporters

36 of 47

Open Standard

37 of 47

Prometheus

38 of 47

Configuration

39 of 47

Prometheus

40 of 47

Ecosystem

41 of 47

Prometheus

42 of 47

https://timber.io/blog/promql-for-humans/

43 of 47

Grafana

44 of 47

Demo

45 of 47

Demo App

  • Spring Boot
  • Three endpoints
  • Chaos events
  • Random walk gauge

46 of 47

Demo App

Grafana

:3000

Prometheus

:9090

Spring Boot

:8080 :8081

PromQL

Scrape

/actuator/prometheus

47 of 47

Questions/Comments?

@Fitzoh