1 of 36

Mastering Experimentation

Deepak Singh

2 of 36

Why should you listen to me?

3 of 36

Mission - Helping PMs create impact in their roles

9k copies sold in Print

10k Subscribers

Growth for PMs

4 of 36

Agenda

  1. Need for experimentation
  2. Necessary conditions for experimentation
  3. Experiments come in all shapes and sizes
  4. Lifecycle of an experiment
  5. Common mistakes
  6. Q&A

Why do we need experimentation?

5 of 36

1

Q: Which experimentation tool does your company have right now?

  • Inhouse
  • VWO
  • Optimizely
  • Don’t have an experimentation program
  • Others

#Assessment

6 of 36

Need for experimentation

Why do we need experimentation?

7 of 36

Need for experimentation

Evidence

8 of 36

Necessary Conditions at Org Level

  • Enough units (users, new users)

Why do we need experimentation?

  • Key metric (north star metric) and success criteria for experimentation program is agreed upon and can be measured accurately
  • Changes are easy to make
    1. Regulation don’t prohibit you
    2. Neither does the HiPPO

9 of 36

Necessary Conditions at Org Level

  • Data-driven

Why do we need experimentation?

  • Willingness to invest in infrastructure
  • Okay with failure

10 of 36

Comfortable with failure?

11 of 36

Experiments come in all shapes and sizes

12 of 36

Google’s 41 shades of Blue / 2009 / +$200 million

13 of 36

Bing longer headline of ads / 2012 / +$100 million

14 of 36

Airbnb - URL level testing for SEO

15 of 36

Airbnb - URL level testing for SEO

16 of 36

So what? Take a toolbox approach as a PM because experiments come in all shapes and sizes

17 of 36

Lifecycle of an Experiment

18 of 36

Start by setting the goal for experimentation

  • For accumulated learnings
  • To avoid goal shifting in-between
  • To create meaningful impact by having focus
  • To ensure cross-functional support and prioritization
  • To get buy-in from leadership

Why do we need to define goal for experimentation?

19 of 36

2

Q: What would be a good goal for an eCommerce platform search team experimentation program?

  1. CTR on search results
  2. User conversion (search → payment)

#Assessment

20 of 36

21 of 36

Start by setting the goal for experimentation

Determine if it’s a server side or client side experiment

Why do we need to define goal for experimentation?

22 of 36

23 of 36

Start by setting the goal for experimentation

Determine if it’s a server side or client side experiment

Determine the sample size

Why do we need to define goal for experimentation?

24 of 36

Sample size depends on p-value

  • p-value is the measure of how much the result is a matter of chance/luck
  • p-value < 0.05 means there is a >95% chance that the result isn’t a matter of chance/luck. 1-p is known as confidence level.

The result has to be statistically significant

25 of 36

Sample size depends on MDE

  • Minimum detectable effect (MDE) is the smallest improvement you want to detect.
  • Suppose you want to test a new drug for dogs that whether it can be used for increasing their height or not. And let’s say there are two cases, where you are saying that:
    • the increase in height should be 1cm(MDE).
    • the increase in height should be 10cm(MDE).

Calculate the sample size to reach significance in advance

26 of 36

3

Q: Where would you need higher sample size?

  • MDE = 1 cm
  • MDE = 10 cm

#Assessment

27 of 36

Why can’t we set MDE as low as possible?

#Assessment

28 of 36

Sample size depends on Power

  • Power is the probability that an experiment will flag a real change as statistically significant.
    • 80% power → 8 of 10 changes detected
    • 60% power → 6 of 10 changes detected

Power

  • Across industries, 80-90% power is considered reasonable

29 of 36

4

Q: What would be the sample size for p=0.01, baseline conversion = 10%, MDE = 10% relative, power = 95%?

  • 25,681
  • 19,826
  • 20,923
  • 32610

#Assessment

30 of 36

Mistakes in Life Cycle of an Experiment

31 of 36

Mistake #1 is not calculating sample size in advance

Calculating the sample size and # of days needed to reach sample size after the experiment starts

Don’t calculate the sample size to reach significance in advance and # of days required

32 of 36

Mistake #2 is not evaluating non-binomial metrics differently

33 of 36

Mistake #3 is peeking

The more the peeking, the higher the chances of false positives as the experiment often dips into significance before coming back out

Don’t calculate the sample size to reach significance in advance and # of days required

34 of 36

Mistake #4 is multiplicity

Multiplicity refers to the “potential inflation of type I error rate (false positive) as a result of multiple testing, for example because of multiple subgroup comparisons, analysis of multiple outcomes, and multiple analyses of the same outcome at different times.”

Don’t calculate the sample size to reach significance in advance and # of days required

35 of 36

36 of 36

Q&A