1 of 45

LESSONS LEARNED

MONITORING THE DATA PIPELINE

2 of 45

AGENDA

  • Who am I?
  • What’s a Hulu?
  • Beacons & the Data Pipeline
  • Monitoring – Take One
  • Monitoring – Take Two

3 of 45

PHOTO WALL

4 of 45

TRISTAN REID

METRICS & REPORTING TOOLS TEAM LEAD

5 of 45

Help people find and enjoy the �world’s premium content –�when, where & how they want it.

HULU MISSION

6 of 45

6

PREMIUM CONTENT

QUALITY AD EXPERIENCE

USER CONTROL

7 of 45

  • Service Oriented
  • Small teams, specialized scopes
  • Build tools for other developers
  • Right tool for the job

7

8 of 45

BEACONS & THE DATA PIPELINE

8

9 of 45

EXTERNAL VIEW OF BEACONS

FIRE & FORGET

HTTP FORMAT

HIGH AVAILABILITY

COLLECT

TRANSFORM

PROCESS

10 of 45

BEACONS

80 2013-04-01 00:00:00

/v3/playback/start?

bitrate=650

&cdn=Akamai

&channel=Anime

&client=Explorer

&computerguid=EA8FA1000232B8F6986C3E0BE55E9333

&contentid=5003673

Which show is the user watching?

Which pages did they visit?

How long did they stay?

Where did they come from?

Did they become Plus members?

11 of 45

THE PIPELINE

BEACON COLLECTION SERVICE

HDFS

HIVE

RDBMS

Log Collector / Flume

MapReduce Jobs

Continuous Aggregation / Selective Publishing

DEVELOPERS

BUSINESS ANALYSTS

MONITORING

REPORTING

12 of 45

DATA COLLECTION�

AVG 12,000 events / second

PEAK ~35K

13 of 45

DATA NEVER STOPS�

AND WE CAN’T LOSE ANY

14 of 45

LOG COLLECTION

DEVICES

CDN

LOAD BALANCER

LOG COLLECTION

machine #1

machine #2

machine #11

machine #3

HDFS

FILES BUCKETED BY BEACON

TYPE AND PARTITIONED BY HOUR

15 of 45

MAPREDUCE - FROM BEACONS TO BASEFACTS

video_id   

289696

content_partner_id    

398

distribution_partner_id   

602

distro_platform_id   

14

is_on_hulu   

0

hourid

383149

watched

76426

16 of 45

HULU MAPREDUCE METRICS JOBS

MapReduce code, including metadata lookups

Beaconspec compiler

Definitions of beacons and base-facts

BeaconSpec DSL

JFlex & CUP

Java (Generated)

Job Scheduler

Scala / Akka

Documentation

Automated Validations for Beacon Generators

IN PROGRESS

17 of 45

USERJOBS

  • Mention the MVEL coolness

MVEL:

client contains 'Chrome' &&

fullscreen == true &&

(os contains 'Windows' || os contains 'Mac') 

18 of 45

AGGREGATION & PUBLISHING

HOURLY FACTS

DAILY / WEEKLY / MONTHLY

QUARTERLY / ANNUAL

AGGREGATIONS

MySQL

SQL

POPULAR DATA

PUBLISHING

19 of 45

REPORTING FLOW

RP2 DB

Published DB’s

Report Controller

Scheduler

Data API Service

HiveRunner

Reporting Portal UI (RP2)

Submit Report

Available Columns / Date Range Checks

Execute Report

Generate Query

Run

Queue

Check

Status

20 of 45

21 of 45

22 of 45

23 of 45

MONITORING

24 of 45

BIG DATA PIPELINE?

BET THAT’S GOING GREAT FOR YOU.

25 of 45

EMAIL EXPLOSIONS

CHANGE

GATEKEEPING

OVERHEAD

CONSUMPTION

26 of 45

LOTS OF MONITORING TOOLS AVAILABLE

Goals – transition sep slide

Cluster

Ingest

Jobs

OpenTSDB & Graphite

C

I

C

J

27 of 45

28 of 45

WHAT’S GOING ON??!??

How is our cluster?

Will we meet our SLAs?

How fast did a job run?

How did runtime compare to historical?

How is this component?

How is our system?

29 of 45

ACCESS ALL YOUR TOOLS IN ONE PLACE...

Comprehensive Web UI

BUT AVOID MULTITASKING

Service Oriented Architecture

30 of 45

31 of 45

32 of 45

33 of 45

34 of 45

DOES THIS SOLVE OUR PROBLEMS?

34

Single point of access?

Maintain services separately?

35 of 45

TAKE THAT!

DATA PIPELINE�ISSUES!

36 of 45

OUR USERS’ PERSPECTIVE?

We detect platform issues

We quickly troubleshoot errors

We track relative performance

We know where we are re: SLAs

…BUT IS DETECTION OF A PROBLEM ENOUGH?

WE NEED TO THINK OF THINGS FROM THE �REPORT USERS’ PERSPECTIVES

37 of 45

THE USER PERSPECTIVE

REPORT USERS

REPORTS

FAILEDSUCCESSFULREPORTS RUN

DATA PIPELINE RESOURCES

ETC!

USER GROUP

USER GROUP

38 of 45

CONTEXTUAL TROUBLESHOOTING MODEL

  • Connect issues to business units
  • Better impact assessment
  • Tune performance per user needs

  • We need a graph data structure, populated with the stuff we care about

  • Something like this

39 of 45

WHY A GRAPH?

…INSTEAD OF RDBMS

  • Indeterminate # of Joins
  • Query for graph connectedness is trivial and short
  • Query for connectedness w/ SQL relies on knowing the intermediate resources

INSTEAD OF A TREE?

  • Data is sometimes recombinant (e.g. a metric in multiple reports to same user)

40 of 45

41 of 45

Let’s investigate…

These failed before getting to a data store

Most of the hive failures were the same table, but it’s a common table

As we filter, the matched reports show up on the bottom of the page. The log link shows us the details

42 of 45

Each service implements a log-fetching interface, specific to the resources used for a particular report

43 of 45

SUCCESS!

44 of 45

IN SUMMARY…

Find the important questions measure the right data

Make troubleshooting easy

Small distinct services are easy to create, maintain, and wire together

45 of 45

QUESTIONS?

THANKS TO…

  • Muthu…the Platform GrandMaster
  • All of Metrics Platform, Tools, Reporting for making this stuff
    • Mohamed, Chris, Charlie, Robert, Phong, AJ, Ratheesh, Adi, Matt, Shashank, Joanne, Siddhartha, Tamir, Jun, James, Dr. Kevin, Hang
  • All of the Hulu DEV team for general awesomeness
  • Prasan…thanks for the impetus to do this. I’ll look u up
  • Kevin…thanks for Hulu. I’ll send u a snap
  • And the awesome Presentation Design team