LESSONS LEARNED
MONITORING THE DATA PIPELINE
AGENDA
PHOTO WALL
TRISTAN REID
METRICS & REPORTING TOOLS TEAM LEAD
Help people find and enjoy the �world’s premium content –�when, where & how they want it.
HULU MISSION
6
PREMIUM CONTENT
QUALITY AD EXPERIENCE
USER CONTROL
7
BEACONS & THE DATA PIPELINE
8
EXTERNAL VIEW OF BEACONS
FIRE & FORGET
HTTP FORMAT
HIGH AVAILABILITY
COLLECT
TRANSFORM
PROCESS
BEACONS
80 2013-04-01 00:00:00
/v3/playback/start?
bitrate=650
&cdn=Akamai
&channel=Anime
&client=Explorer
&computerguid=EA8FA1000232B8F6986C3E0BE55E9333
&contentid=5003673
…
Which show is the user watching?
Which pages did they visit?
How long did they stay?
Where did they come from?
Did they become Plus members?
THE PIPELINE
BEACON COLLECTION SERVICE
HDFS
HIVE
RDBMS
Log Collector / Flume
MapReduce Jobs
Continuous Aggregation / Selective Publishing
DEVELOPERS
BUSINESS ANALYSTS
MONITORING
REPORTING
DATA COLLECTION�
AVG 12,000 events / second
PEAK ~35K
DATA NEVER STOPS�
AND WE CAN’T LOSE ANY
LOG COLLECTION
DEVICES
CDN
LOAD BALANCER
LOG COLLECTION
machine #1
machine #2
…
machine #11
machine #3
HDFS
FILES BUCKETED BY BEACON
TYPE AND PARTITIONED BY HOUR
MAPREDUCE - FROM BEACONS TO BASEFACTS
| |
video_id | 289696 |
content_partner_id | 398 |
distribution_partner_id | 602 |
distro_platform_id | 14 |
is_on_hulu | 0 |
… | |
hourid | 383149 |
watched | 76426 |
HULU MAPREDUCE METRICS JOBS
MapReduce code, including metadata lookups
Beaconspec compiler
Definitions of beacons and base-facts
BeaconSpec DSL
JFlex & CUP
Java (Generated)
Job Scheduler
Scala / Akka
Documentation
Automated Validations for Beacon Generators
IN PROGRESS
USERJOBS
MVEL:
client contains 'Chrome' &&
fullscreen == true &&
(os contains 'Windows' || os contains 'Mac')
AGGREGATION & PUBLISHING
HOURLY FACTS
DAILY / WEEKLY / MONTHLY
QUARTERLY / ANNUAL
AGGREGATIONS
MySQL
SQL
POPULAR DATA
PUBLISHING
REPORTING FLOW
RP2 DB
Published DB’s
Report Controller
Scheduler
Data API Service
HiveRunner
Reporting Portal UI (RP2)
Submit Report
Available Columns / Date Range Checks
Execute Report
Generate Query
Run
Queue
Check
Status
MONITORING
BIG DATA PIPELINE?
BET THAT’S GOING GREAT FOR YOU.
EMAIL EXPLOSIONS
CHANGE
GATEKEEPING
OVERHEAD
CONSUMPTION
LOTS OF MONITORING TOOLS AVAILABLE
Goals – transition sep slide
Cluster
Ingest
Jobs
OpenTSDB & Graphite
C
I
C
J
WHAT’S GOING ON??!??
How is our cluster?
Will we meet our SLAs?
How fast did a job run?
How did runtime compare to historical?
How is this component?
How is our system?
ACCESS ALL YOUR TOOLS IN ONE PLACE...
Comprehensive Web UI
BUT AVOID MULTITASKING
Service Oriented Architecture
DOES THIS SOLVE OUR PROBLEMS?
34
Single point of access?
Maintain services separately?
TAKE THAT!
DATA PIPELINE�ISSUES!
OUR USERS’ PERSPECTIVE?
We detect platform issues
We quickly troubleshoot errors
We track relative performance
We know where we are re: SLAs
…BUT IS DETECTION OF A PROBLEM ENOUGH?
WE NEED TO THINK OF THINGS FROM THE �REPORT USERS’ PERSPECTIVES
THE USER PERSPECTIVE
REPORT USERS
REPORTS
FAILED �SUCCESSFUL�REPORTS RUN
DATA PIPELINE RESOURCES
ETC!
USER GROUP
USER GROUP
CONTEXTUAL TROUBLESHOOTING MODEL
WHY A GRAPH?
…INSTEAD OF RDBMS
…INSTEAD OF A TREE?
Let’s investigate…
These failed before getting to a data store
Most of the hive failures were the same table, but it’s a common table
As we filter, the matched reports show up on the bottom of the page. The log link shows us the details
Each service implements a log-fetching interface, specific to the resources used for a particular report
SUCCESS!
IN SUMMARY…
Find the important questions measure the right data
Make troubleshooting easy
Small distinct services are easy to create, maintain, and wire together
QUESTIONS?
THANKS TO…