1 of 26

A/B Tests with Unobserved Network Spillover: Design and Inference

JON STALLRICH, NIC LARSEN, SRIJAN SENGUPTA

NC STATE UNIVERSITY

DAE 2024

2 of 26

Online controlled experiments (OCEs)

1

    • Online businesses (e.g., Facebook, Netflix, and Google) deploy thousands of OCEs a year on tens of millions of users

    • Google’s 41 Shades of Blue experiment: $200M profit

    • Bing’s small change in Ad text: $100M profit

3 of 26

Statistical challenges in A/B testing

2

    • Focus on A/B testing: control vs test treatment

    • Sensitivity and Small Treatment Effects
    • Heterogenous Treatment Effects
    • Short-Term vs Long-Term Treatment Effects
    • Optional Stopping (i.e., sequential testing)
    • Interference

4 of 26

General framework

3

 

5 of 26

Average treatment effect

4

    • Target of inference

    • Problem 1: experiment can’t apply either treatment globally

    • Problem 2: only observe one response per unit

    • Assumptions needed for inference to be possible

 

6 of 26

SUTVA and Difference of Means

5

 

 

7 of 26

Interference violates SUTVA

6

    • Example: LinkedIn studies new feature to increase total messages sent

Unit’s messaging behavior

interferes with messaging behavior of their connected users (neighbors)

8 of 26

Cluster-based randomization

7

    • Identify connections between users (network) and assign treatments to clusters of neighbors

    • Gui et al (2015), Saint-Jacques (2019), Larsen et al (2023)
    • Pro: Reduces interference between test and control
    • Con: Assumes perfect knowledge of network

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

Adjacency Matrix

9 of 26

Optimal design approach

8

 

10 of 26

Interference Model with Covariate

9

 

11 of 26

Expected value of difference of means

10

 

 

12 of 26

 

11

 

 

13 of 26

Cluster-based design

12

 

 

1

2

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

14 of 26

Potential issues with cluster-based

13

 

15 of 26

HODOR: Hold-out Design for Online Randomized Experiments

14

 

+

1

16 of 26

HODOR Estimator

15

 

17 of 26

HODOR: Optimal design

16

 

 

 

 

 

 

18 of 26

Simulation study: Absolute error

17

 

Difference of means

19 of 26

Simulation study: Absolute error

18

 

Strong

Mild

None

20 of 26

Summary

19

 

21 of 26

References

    • Gui, H., Xu, Y., Bhasin, A., and Han, J. (2015). Network a/b testing: From sampling toestimation. In Proceedings of the 24th International Conference on World Wide Web, WWW’15, pages 399–409. International World Wide Web Conferences Steering Committee.
    • Saint-Jacques, G., Varshney, M., Simpson, J., and Xu, Y. (2019). Using ego-clusters to measure network effects at linkedin. arXiv preprint arXiv:1903.08755.
    • Larsen, N., Stallrich, J., Sengupta, S., Deng, A., Kohavi, R., and Stevens, N. T. (2023) Statistical Challenges in Online Controlled Experiments: A Review of A/B Testing Methodology, The American Statistician
    • Parker, B. M., Gilmour, S. G., and Schormans, J. (2017). Optimal design of experiments on connected units with application to social networks. Journal of the Royal Statistical Society. Series C (Applied Statistics), pages 455–480.
    • Basse, G. W. and Airoldi, E. M. (2018). “Model-assisted design of experiments in the presence of network-correlated outcomes”. Biometrika, 105(4):849–858. 
    • Pokhilko, V., Zhang, Q., and Kang, L. (2019). “D-optimal design for network a/b testing”. Journal of Statistical Theory and Practice, 13(4):1–23. 
    • Zhang, Q. and Kang, L. (2022). “Locally optimal design for a/b tests in the presence of covariates and network dependence”. Technometrics, 64(3):358–369.

20

22 of 26

 

21

 

23 of 26

 

22

 

24 of 26

 

23

 

25 of 26

HODOR: Randomization inference

24

 

26 of 26

Simulation Study 2: Power and coverage

25