Mastering Experimentation
Deepak Singh
Why should you listen to me?
Mission - Helping PMs create impact in their roles
9k copies sold in Print
10k Subscribers
Growth for PMs
Agenda
Why do we need experimentation?
1
Q: Which experimentation tool does your company have right now?
#Assessment
Need for experimentation
Why do we need experimentation?
Need for experimentation
Evidence
Necessary Conditions at Org Level
Why do we need experimentation?
Necessary Conditions at Org Level
Why do we need experimentation?
Comfortable with failure?
Experiments come in all shapes and sizes
Google’s 41 shades of Blue / 2009 / +$200 million
Bing longer headline of ads / 2012 / +$100 million
Airbnb - URL level testing for SEO
Airbnb - URL level testing for SEO
So what? Take a toolbox approach as a PM because experiments come in all shapes and sizes
Lifecycle of an Experiment
Start by setting the goal for experimentation
Why do we need to define goal for experimentation?
2
Q: What would be a good goal for an eCommerce platform search team experimentation program?
#Assessment
Start by setting the goal for experimentation
Determine if it’s a server side or client side experiment
Why do we need to define goal for experimentation?
Start by setting the goal for experimentation
Determine if it’s a server side or client side experiment
Determine the sample size
Why do we need to define goal for experimentation?
Sample size depends on p-value
The result has to be statistically significant
Sample size depends on MDE
Calculate the sample size to reach significance in advance
3
Q: Where would you need higher sample size?
#Assessment
Why can’t we set MDE as low as possible?
#Assessment
Sample size depends on Power
Power
4
Q: What would be the sample size for p=0.01, baseline conversion = 10%, MDE = 10% relative, power = 95%?
#Assessment
Mistakes in Life Cycle of an Experiment
Mistake #1 is not calculating sample size in advance
Calculating the sample size and # of days needed to reach sample size after the experiment starts
Don’t calculate the sample size to reach significance in advance and # of days required
Mistake #2 is not evaluating non-binomial metrics differently
Mistake #3 is peeking
The more the peeking, the higher the chances of false positives as the experiment often dips into significance before coming back out
Don’t calculate the sample size to reach significance in advance and # of days required
Mistake #4 is multiplicity
Multiplicity refers to the “potential inflation of type I error rate (false positive) as a result of multiple testing, for example because of multiple subgroup comparisons, analysis of multiple outcomes, and multiple analyses of the same outcome at different times.”
Don’t calculate the sample size to reach significance in advance and # of days required
Q&A