PUBLIC

Lighthouse Metric Variability and Accuracy

Author: phulce@chromium.org 
Status: complete

Last Updated: January 5, 2018

Executive Summary

Variance (least to most): Lantern < DevTools throttled < WPT throttled < Unthrottled

Example variance impact in DevTools throttled for a site with a performance score of 50:

95% confidence interval for 1-run: 50 +/- 15  (score range 65-35)
95% confidence interval for 3-runs: 50 +/- 11 (score range 61-39)
95% confidence interval for 5-runs: 50 +/- 8 (score range 58-42)

Accuracy (most to least): WPT throttled > DevTools throttled > Lantern > Unthrottled

Total combined single-run error[1] (systematic inaccuracy + variance) for Lantern: +/- 45%
Total combined single-run error (systematic inaccuracy + variance) for
DevTools: +/- 47%
Total combined single-run error (systematic inaccuracy + variance) for
WPT: +/- 15%

DevTools correlates with WPT at ~.89
Lantern
 correlates with WPT at ~.87
Lantern
 correlates with DevTools at ~.87

Lighthouse Metric Variability and Accuracy

Executive Summary

Glossary

Performance Metric Goals

Variability

Sources of Variability

Local Network Variability

Tier-1 Network Variability

Web Server Variability

Client Hardware Variability

Client Resource Contention

Browser Nondeterminism

Page Nondeterminism

Analysis

First Contentful Paint

First Meaningful Paint

Time to Consistently Interactive

Notes and Conclusions

Approaches with the Lowest Variance

"Accuracy" of 1/3/5 Runs

1-Run Reference Stats

3-Run Reference Stats

5-Run Reference Stats

Re-evaluating Previous Accuracy Data with Variability Data

First Attempt - Lantern vs. DT Throttling

Second Attempt - Lantern vs. WPT

Second Attempt - DT Throttling vs. WPT

Accuracy

Selecting a Target

Analysis

First Contentful Paint

First Meaningful Paint

Time to Consistently Interactive

Notes and Conclusions

First Contentful Paint

First Meaningful Paint

Time to Consistently Interactive

Test Environment Settings

LightRider

Variance

1-Run Reference Stats

3-Run Reference Stats

Accuracy

Conclusions

Future Work

Related Documents

Glossary

  • DT/DevTools - Refers to Chrome DevTools. In the context of this variability and accuracy document, it refers to the current Lighthouse implementation in DevTools that uses Chrome's built-in network and CPU throttling mechanisms to slow down page load and report the observed metric timestamps.
  • Lantern - Refers to Project Lantern in Lighthouse that speeds up run time by loading the page as fast as possible and simulating on its own what the time under 3G speeds would be. This project was a requirement of the LightRider effort where throttling will not be available.
  • WPT - Refers to WebPageTest.org where Lighthouse can be run on real mobile devices with high-fidelity network throttling, as close to a real user experience as possible.
  • FCP - Refers to the metric First Contentful Paint, the first time an image or text appears on screen.
  • FMP - Refers to the metric First Meaningful Paint, the first time a majority of above-the-fold elements appear on screen.
  • TTI - Refers to the metric Time to Consistently Interactive, the first time after network quiet that there is a prolonged period of CPU inactivity.

Performance Metric Goals

Lighthouse performance scoring has two primary goals which we strive to achieve.

  1. Performance scores should reflect real-world user experiences.
  2. Performance scores should be reproducible.

Lighthouse's approach to meet these two goals is through the assessment of performance metrics. Progressive web metrics (FCP, FMP, TTI) are designed to capture the user experience and are the sole input to the Lighthouse performance score. To measure our success on these goals, we establish the following subgoals.

  1. A page ranked in the Xth percentile according to real-world environment should be ranked in the Xth percentile according to Lighthouse.
  2. A page ranked in the Xth percentile according to Lighthouse on a given run should be ranked in the Xth percentile when run a second time.

These two subgoals track, correspondingly, the accuracy and variability of the metrics.

Variability

While accuracy is first and foremost important, it cannot be examined without understanding the inherent variability involved with performance measurement. This section discusses the different sources of variability and contains analysis that quantifies the different mitigations.

Sources of Variability

Variability in performance measurement is introduced via a number of channels with different levels of impact. Below is a table containing several common sources of metric variability, the typical impact they have on results, and the extent to which different Lighthouse runtimes are able to mitigate their effect.

Source

Impact

Lantern

DT Throttling

WPT Throttling

Local network variability

High

MITIGATED

PARTIALLY MITIGATED

PARTIALLY MITIGATED

Tier-1 network variability

Medium

MITIGATED

PARTIALLY MITIGATED

PARTIALLY MITIGATED

Web server variability

Low

NO MITIGATION

PARTIALLY MITIGATED

NO MITIGATION

Client hardware variability

High

PARTIALLY MITIGATED

NO MITIGATION

NO MITIGATION

Client resource contention

High

PARTIALLY MITIGATED

NO MITIGATION

NO MITIGATION

Browser nondeterminism

Medium

PARTIALLY MITIGATED

NO MITIGATION

NO MITIGATION

Page nondeterminism

Medium

NO MITIGATION

NO MITIGATION

NO MITIGATION

Below is a table containing several common sources of metric variability, the typical impact they have on results, and the extent to which they are likely to occur in different environments.

Source

Impact

Typical End User

LightRider

Controlled Lab

Local network variability

High

LIKELY

UNLIKELY

UNLIKELY

Tier-1 network variability

Medium

POSSIBLE

POSSIBLE

POSSIBLE

Web server variability

Low

LIKELY

LIKELY

LIKELY

Client hardware variability

High

LIKELY

UNLIKELY

UNLIKELY

Client resource contention

High

LIKELY

POSSIBLE

UNLIKELY

Browser nondeterminism

Medium

CERTAIN

CERTAIN

CERTAIN

Page nondeterminism

Medium

LIKELY

LIKELY

LIKELY

Below are more detailed descriptions of the sources of variance and the impact they have on the most likely combinations of Lighthouse runtime + environment. While DevTools throttling and Lantern approaches could be used in any of these three environments, the typical end user exclusively uses DevTools throttling on their machine and Lantern was designed to be able to use in LightRider. WPT, on the other hand, is always a controlled lab setting.

Local Network Variability

Local networks have inherent variability from packet loss, variable traffic prioritization, and last-mile network congestion. Users with cheap routers and many devices sharing limited bandwidth are usually the most susceptible to this. DevTools partially mitigates these effects by applying a minimum request latency and maximum throughput that masks underlying retries. Lantern mitigates these effects by replaying network activity on its own. WPT mitigates these effects with additional packet latency, maximum throughput, and a high quality underlying network of a lab environment.

Tier-1 Network Variability

Network interconnects are generally very stable and have minimal impact but cross-geo requests, i.e. measuring performance of a Chinese site from the US, can start to experience a high degree of latency introduced from tier-1 network hops. DevTools and WPT partially mask these effects with network throttling. Lantern mitigates these effects by replaying network activity on its own.

Web Server Variability

Web servers have variable load and do not always respond with the same delay. Lower traffic sites with shared hosting infrastructure are typically more susceptible to this. DevTools partially masks these effects by applying a minimum request latency in its network throttling. Lantern and WPT are both susceptible to this effect but the overall impact is usually low when compared to other network variability.

Client Hardware Variability

The hardware on which the webpage is loading can greatly impact performance. DevTools cannot do much to mitigate this issue. Lantern partially mitigates this issue by capping the theoretical execution time of CPU tasks during simulation. WPT mitigates this issue with a lab environment by maintaining a fleet of identical hardware.

Client Resource Contention

Other applications running on the same machine while Lighthouse is running can cause contention for CPU, memory, and network resources. Malware, browser extensions, and anti-virus software have particularly strong impacts on web performance. Multi-tenant or parallel server environments can also suffer from these issues. DevTools is susceptible to this issue. Lantern partially mitigates this issue by replaying network activity on its own and capping CPU execution. WPT mitigates this issue with lab environment by dedicating each mobile device to a single test and eliminating all other software that might typically interfere with a user's run.

Browser Nondeterminism

Browsers have inherent variability in their execution of tasks that impacts the way webpages are loaded. This is unavoidable for DevTools and WPT as at the end of the day they are simply reporting whatever was observed by the browser. Lantern is able to partially mitigate this effect by simulating execution on its own and only re-using task execution times from the browser in its estimate.

Page Nondeterminism

Pages can contain logic that is nondeterministic that changes the way a user experiences a page, i.e. an A/B test that changes the layout and assets loaded or a different ad experience based on campaign progress. This is an intentional and irremovable source of variance. If the page changes in a way that hurts performance, Lighthouse should be able to identify this case. The only mitigation here is on the part of the site owner in ensuring that the exact same version of the page is being tested between different runs.

Analysis

For the purposes of Lighthouse, we care most about sources of variability which are not easily solved by running in a lab environment. To quantify the variance and assess the success of our mitigations, all following data was collected in a controlled setting meaning that remaining variance is mostly a result of the sources below.

  • Tier-1 Network Variability
  • Web Server Variability
  • Browser nondeterminism
  • Page nondeterminism

300 URLs—a subset of Alexa Top 1000, HTTPArchive URLs, and ad landing pages—were tested 9 times with each of our 4 approaches: DevTools Throttling, Lantern, and WPT Throttling plus a control. To quantify variance, we introduce 5 measures of success on each of our 3 headline metrics FCP, FMP, and TTI.

  • Average Ratio of Standard Deviation to Median
  • Percent of observations that were within 5% of the median
  • Percent of observations that were within 10% of the median
  • Percent of observations that were within 25% of the median
  • Percent of observations that were within 50% of the median

The results for each approach on each metric can be found in the tables below.

First Contentful Paint

Measure

Lantern

DT Throttling

WPT Throttling

Control

Std. / Median

8.0%

9.8%

13.3%

16.8%

Obvs. within 5%

77.18%

79.36%

58.13%

50.73%

Obvs. within 10%

88.46%

89.81%

78.68%

70.32%

Obvs. within 25%

95.76%

96.10%

93.57%

89.00%

Obvs. within 50%

98.43%

98.39%

97.01%

95.78%

First Meaningful Paint

Measure

Lantern

DT Throttling

WPT Throttling

Control

Std. / Median

8.4%

10.6%

13.6%

16.2%

Obvs. within 5%

72.71%

77.59%

57.87%

50.39%

Obvs. within 10%

85.74%

88.50%

77.57%

71.27%

Obvs. within 25%

95.26%

94.79%

92.81%

89.82%

Obvs. within 50%

98.55%

97.69%

97.06%

95.81%

Time to Consistently Interactive

Measure

Lantern

DT Throttling

WPT Throttling

Control

Std. / Median

7.3%

14.1%

14.0%

18.9%

Obvs. within 5%

71.60%

63.68%

52.00%

47.26%

Obvs. within 10%

87.35%

76.67%

72.00%

66.58%

Obvs. within 25%

96.94%

90.50%

89.97%

86.20%

Obvs. within 50%

99.04%

96.36%

97.52%

94.94%

Notes and Conclusions

Approaches with the Lowest Variance

The final ranked list of approaches in increasing order of variance is Lantern, DevTools Throttling, WPT Throttling, and finally the control group (unthrottled). This matches the expectations we have from the number of sources of variance each approach mitigates.

Lantern eliminates almost all network variance and thus achieves lowest overall variance. DevTools throttling applies at the request level rather than packet level and essentially floors all network requests to a minimum of 600 ms which is quite effective at masking underlying variance, at the expense of accuracy. WebPageTest applies additive latency to each packet which somewhat masks variance just by reducing the percentage effect it has on outcomes. As expected, an unthrottled view of the page has the most observed variance as no mitigating factors are in play.

"Accuracy" of 1/3/5 Runs

To assess the accuracy of single observations and quantify how much lower the variance becomes with more observations, we can compute the same success metrics on different random subsets of the data. Below are the 95th percentile and average spearman's rho/MAPE values of our other observation-based approaches for 1, 3, and 5 runs--that is to say, these are the values of our success metrics we might expect to see when using n-observations of a metric.

The tables can be found below. In summary, across all metrics, variance and MAPE decreases by roughly 25% when moving from 1 run to 3 runs and by roughly 50% when moving from 1 run to 5 runs. A 95% confidence interval on the performance score of a site with a median FMP for a 1-run case would span 71-33, 3-run case would span 62-39, 5-run case would span 58-42. A site with a median TTI for a 1-run case would span 65-38, 3-run case would span 61-40, 5-run case would span 58-43.

1-Run Reference Stats

Metric + Approach

Worst Case

95th Percentile

Average

Std. / Mean

FCP - DevTools

.806 - 34.3%

.945 - 12.6%

.961 - 8.48%

9.2%

FCP - WPT

.678 - 49.7%

.891 - 20.7%

.921 - 14.6%

13.2%

FCP - Unthrottled

.710 - 52.5%

.916 - 22.6%

.939 - 16.9%

16.3%

FMP - DevTools

.760 - 36.2%

.926 - 14.6%

.943 - 9.77%

10.0%

FMP - WPT

.669 - 51.4%

.888 - 21.3%

.918 - 15.0%

13.6%

FMP - Unthrottled

.704 - 50.8%

.921 - 19.2%

.938 - 16.3%

15.9%

TTCI - DevTools

.733 - 49.0%

.907 - 17.9%

.932 - 14.3%

13.2%

TTCI - WPT

.779 - 44.0%

.929 - 16.8%

.948 - 14.5%

13.9%

TTCI - Unthrottled

.608 - 66.0%

.884 - 23.8%

.911 - 19.9%

18.3%

3-Run Reference Stats

Metric + Approach

95th Percentile

Std. / Mean

FCP - DevTools

.979 - 7.9%

7.3%

FCP - WPT

.955 - 13.1%

10.6%

FCP - Unthrottled

.970 - 11.4%

12.7%

FMP - DevTools

.965 - 9.2%

8.0%

FMP - WPT

.951 - 11.8%

10.2%

FMP - Unthrottled

.965 - 11.8%

12.2%

TTCI - DevTools

.958 - 8.1%

9.8%

TTCI - WPT

.971 - 9.4%

10.7%

TTCI - Unthrottled

.954 - 12.9%

14.1%

5-Run Reference Stats

Metric + Approach

95th Percentile

Std. / Mean

FCP - DevTools

.985 - 4.9%

4.9%

FCP - WPT

.957 - 12.3%

8.4%

FCP - Unthrottled

.978 - 8.6%

10.1%

FMP - DevTools

.975 - 5.3%

5.0%

FMP - WPT

.956 - 9.8%

8.4%

FMP - Unthrottled

.980 - 9.5%

9.7%

TTCI - DevTools

.969 - 6.3%

7.8%

TTCI - WPT

.975 - 7.5%

8.6%

TTCI - Unthrottled

.970 - 9.4%

11.0%

Re-evaluating Previous Accuracy Data with Variability Data

Previous attempts to assess the accuracy of Lantern have been limited by the fact that only observations from a single run have been used. As a result, it was not possible to tell if Lantern was getting more/less accurate or we were just getting more/less lucky. Now with variability data in hand, we can go back and analyze how likely these outcomes would be if Lantern were 100% accurate and test our null hypothesis using the 1-run reference stats above.

First Attempt - Lantern vs. DT Throttling

The first attempt to assess Lantern accuracy found a spearman's rho of ~.85 and MAPE of ~20% for FCP/FMP and spearman's rho of ~.90 and MAPE ~27% for TTCI. Given our 1-run reference stats, we can safely reject the null hypothesis for these outcomes and observe an inaccuracy inherent to Lantern of ~6-10%.

Second Attempt - Lantern vs. WPT

The second attempt to assess Lantern accuracy found a spearman's rho of ~.78 and MAPE of ~32% for FCP/FMP and spearman's rho of ~.88 and MAPE ~33% for TTCI. However, when attempting to replicate those particular WPT results future WPT runs observed similarly high error, in fact even higher than that of Lantern. Given our 1-run reference stats, we would normally reject the null hypothesis for these outcomes and observe an inaccuracy inherent to Lantern of ~5-10%, but the inability to replicate suggests that this particular run was flawed and this stat cannot necessarily be trusted.  

Second Attempt - DT Throttling vs. WPT

The second attempt also assessed DevTools throttling accuracy and found a spearman's rho of ~.82 and MAPE of ~30% for FCP/FMP and spearman's rho of ~.82 and MAPE ~40% for TTCI. However, when attempting to replicate those particular WPT results future WPT runs observed similarly high error, in fact even higher than that of DevTools throttling. Given our 1-run reference stats, we would normally reject the null hypothesis for these outcomes and observe an inaccuracy inherent to DevTools throttling of at least ~6-10%, but the inability to replicate suggests that this particular run was flawed and this stat cannot necessarily be trusted.

Accuracy

Accuracy is our primary objective, and this section explores how Lighthouse should measure accuracy as well as the accuracy of various Lighthouse throttling approaches. When acting on information from a single observation of a metric, its accuracy will be a combination of the metric's accuracy on average and its variance. For example, a single observation of a metric whose mean is 100% accurate but rarely reports within 5% of that mean shouldn't be trusted more than a metric whose mean is 5% off and has no variance.

Selecting a Target

Lighthouse performance recommendations are designed for the mobile web with our prototypical user chosen on a Nexus 5X on a 3G cellular connection (1.6Mbps down, 750kbps up, 150ms RTT). Ideally, we would examine a corpus of data from identical user devices and connections. Unfortunately, no such dataset exists as the Chrome UX report and UKM data is limited in its ability to distinguish identical connection characteristics, instead stopping at the granularity of effectiveConnectionType. Luckily, one of our approaches allows us to create this dataset. In fact, WebPageTest with real phones and high fidelity network throttling is almost exactly the user profile we are targeting. Thus, Lighthouse with WPT throttling will be our gold standard for accuracy. Note that this means by definition WPT will have a systematic error of 0.

Analysis

Tables below contain our two primary accuracy measures—spearman's rho and MAPE—for the matrix of FCP/FMP/TTCI between Unthrottled, Lantern, DevTools, and WPT. These statistics are computed using the same 300 URL, 9-run dataset as above and reflect how closely the median values track that of a different Lighthouse throttling approach.

First Contentful Paint

DevTools

WPT

Unthrottled

.738 - 27.1%

.691 - 33.8%

Lantern

.811 - 23.1%

.785 - 28.3%

DevTools

--

.855 - 22.3%

First Meaningful Paint

DevTools

WPT

Unthrottled

.694 - 33.8%

.635 - 40.5%

Lantern

.811 - 23.6%

.761 - 33.7%

DevTools

--

.813 - 27.0%

Time to Consistently Interactive

DevTools

WPT

Unthrottled

.743 - 62.0%

.712 - 66.4%

Lantern

.869 - 42.5%

.854 - 45.4%

DevTools

--

.889 - 32.3%

Notes and Conclusions

We can combine this analysis of systematic error with the earlier variance measurements to arrive at conclusions for how much error we expect single observations in each environment to contain. Note that WPT will have a clear advantage here as only its variance contributes to total combined error.

We conclude that a single observation from DevTools can be expected to be ~30-45% off from true 3G performance and Lantern is consistently ~6% more inaccurate than DevTools.

First Contentful Paint

Systematic Error

Variance

Total

Lantern

28.3%

8.0%

36.3%

DevTools

22.3%

8.5%

30.8%

WPT

0%

14.6%

14.6%

Unthrottled

33.8%

16.9%

50.7%

First Meaningful Paint

Systematic Error

Variance

Total

Lantern

33.7%

8.4%

42.1%

DevTools

27.0%

9.8%

36.8%

WPT

0%

15.0%

15.0%

Unthrottled

40.5%

16.3%

56.8%

Time to Consistently Interactive

Systematic Error

Variance

Total

Lantern

45.4%

7.3%

52.7%

DevTools

32.3%

14.3%

46.6%

WPT

0%

14.5%

14.5%

Unthrottled

66.4%

19.9%

86.3%

Test Environment Settings

  • Lantern and unthrottled traces were collected on an HP Z840 Workstation using the Lighthouse CLI with mobile browser emulation on and all throttling disabled.
  • DevTools throttled traces were collected on an HP Z840 Workstation using the Lighthouse CLI with mobile browser emulation on, 4x CPU throttling, and DevTools "Fast 3G" network throttling.
  • WebPageTest throttled traces were collected on Moto G (gen 4) devices in Dulles location with "3G Fast" network throttling.

[1] Sum of MAPE due to variance of the environment from single-run to single-run and MAPE due to predicting 3G performance on WPT using multiple runs. Note that in the case of WPT, systematic error will be 0.