Author: phulce@chromium.org
Status: complete
Last Updated: January 5, 2018
Variance (least to most): Lantern < DevTools throttled < WPT throttled < Unthrottled
Example variance impact in DevTools throttled for a site with a performance score of 50:
95% confidence interval for 1-run: 50 +/- 15 (score range 65-35)
95% confidence interval for 3-runs: 50 +/- 11 (score range 61-39)
95% confidence interval for 5-runs: 50 +/- 8 (score range 58-42)
Accuracy (most to least): WPT throttled > DevTools throttled > Lantern > Unthrottled
Total combined single-run error[1] (systematic inaccuracy + variance) for Lantern: +/- 45%
Total combined single-run error (systematic inaccuracy + variance) for DevTools: +/- 47%
Total combined single-run error (systematic inaccuracy + variance) for WPT: +/- 15%
DevTools correlates with WPT at ~.89
Lantern correlates with WPT at ~.87
Lantern correlates with DevTools at ~.87
Lighthouse Metric Variability and Accuracy
Time to Consistently Interactive
Approaches with the Lowest Variance
Re-evaluating Previous Accuracy Data with Variability Data
First Attempt - Lantern vs. DT Throttling
Second Attempt - Lantern vs. WPT
Second Attempt - DT Throttling vs. WPT
Time to Consistently Interactive
Time to Consistently Interactive
Lighthouse performance scoring has two primary goals which we strive to achieve.
Lighthouse's approach to meet these two goals is through the assessment of performance metrics. Progressive web metrics (FCP, FMP, TTI) are designed to capture the user experience and are the sole input to the Lighthouse performance score. To measure our success on these goals, we establish the following subgoals.
These two subgoals track, correspondingly, the accuracy and variability of the metrics.
While accuracy is first and foremost important, it cannot be examined without understanding the inherent variability involved with performance measurement. This section discusses the different sources of variability and contains analysis that quantifies the different mitigations.
Variability in performance measurement is introduced via a number of channels with different levels of impact. Below is a table containing several common sources of metric variability, the typical impact they have on results, and the extent to which different Lighthouse runtimes are able to mitigate their effect.
Source | Impact | Lantern | DT Throttling | WPT Throttling |
Local network variability | High | MITIGATED | PARTIALLY MITIGATED | PARTIALLY MITIGATED |
Tier-1 network variability | Medium | MITIGATED | PARTIALLY MITIGATED | PARTIALLY MITIGATED |
Web server variability | Low | NO MITIGATION | PARTIALLY MITIGATED | NO MITIGATION |
Client hardware variability | High | PARTIALLY MITIGATED | NO MITIGATION | NO MITIGATION |
Client resource contention | High | PARTIALLY MITIGATED | NO MITIGATION | NO MITIGATION |
Browser nondeterminism | Medium | PARTIALLY MITIGATED | NO MITIGATION | NO MITIGATION |
Page nondeterminism | Medium | NO MITIGATION | NO MITIGATION | NO MITIGATION |
Below is a table containing several common sources of metric variability, the typical impact they have on results, and the extent to which they are likely to occur in different environments.
Source | Impact | Typical End User | LightRider | Controlled Lab |
Local network variability | High | LIKELY | UNLIKELY | UNLIKELY |
Tier-1 network variability | Medium | POSSIBLE | POSSIBLE | POSSIBLE |
Web server variability | Low | LIKELY | LIKELY | LIKELY |
Client hardware variability | High | LIKELY | UNLIKELY | UNLIKELY |
Client resource contention | High | LIKELY | POSSIBLE | UNLIKELY |
Browser nondeterminism | Medium | CERTAIN | CERTAIN | CERTAIN |
Page nondeterminism | Medium | LIKELY | LIKELY | LIKELY |
Below are more detailed descriptions of the sources of variance and the impact they have on the most likely combinations of Lighthouse runtime + environment. While DevTools throttling and Lantern approaches could be used in any of these three environments, the typical end user exclusively uses DevTools throttling on their machine and Lantern was designed to be able to use in LightRider. WPT, on the other hand, is always a controlled lab setting.
Local networks have inherent variability from packet loss, variable traffic prioritization, and last-mile network congestion. Users with cheap routers and many devices sharing limited bandwidth are usually the most susceptible to this. DevTools partially mitigates these effects by applying a minimum request latency and maximum throughput that masks underlying retries. Lantern mitigates these effects by replaying network activity on its own. WPT mitigates these effects with additional packet latency, maximum throughput, and a high quality underlying network of a lab environment.
Network interconnects are generally very stable and have minimal impact but cross-geo requests, i.e. measuring performance of a Chinese site from the US, can start to experience a high degree of latency introduced from tier-1 network hops. DevTools and WPT partially mask these effects with network throttling. Lantern mitigates these effects by replaying network activity on its own.
Web servers have variable load and do not always respond with the same delay. Lower traffic sites with shared hosting infrastructure are typically more susceptible to this. DevTools partially masks these effects by applying a minimum request latency in its network throttling. Lantern and WPT are both susceptible to this effect but the overall impact is usually low when compared to other network variability.
The hardware on which the webpage is loading can greatly impact performance. DevTools cannot do much to mitigate this issue. Lantern partially mitigates this issue by capping the theoretical execution time of CPU tasks during simulation. WPT mitigates this issue with a lab environment by maintaining a fleet of identical hardware.
Other applications running on the same machine while Lighthouse is running can cause contention for CPU, memory, and network resources. Malware, browser extensions, and anti-virus software have particularly strong impacts on web performance. Multi-tenant or parallel server environments can also suffer from these issues. DevTools is susceptible to this issue. Lantern partially mitigates this issue by replaying network activity on its own and capping CPU execution. WPT mitigates this issue with lab environment by dedicating each mobile device to a single test and eliminating all other software that might typically interfere with a user's run.
Browsers have inherent variability in their execution of tasks that impacts the way webpages are loaded. This is unavoidable for DevTools and WPT as at the end of the day they are simply reporting whatever was observed by the browser. Lantern is able to partially mitigate this effect by simulating execution on its own and only re-using task execution times from the browser in its estimate.
Pages can contain logic that is nondeterministic that changes the way a user experiences a page, i.e. an A/B test that changes the layout and assets loaded or a different ad experience based on campaign progress. This is an intentional and irremovable source of variance. If the page changes in a way that hurts performance, Lighthouse should be able to identify this case. The only mitigation here is on the part of the site owner in ensuring that the exact same version of the page is being tested between different runs.
For the purposes of Lighthouse, we care most about sources of variability which are not easily solved by running in a lab environment. To quantify the variance and assess the success of our mitigations, all following data was collected in a controlled setting meaning that remaining variance is mostly a result of the sources below.
300 URLs—a subset of Alexa Top 1000, HTTPArchive URLs, and ad landing pages—were tested 9 times with each of our 4 approaches: DevTools Throttling, Lantern, and WPT Throttling plus a control. To quantify variance, we introduce 5 measures of success on each of our 3 headline metrics FCP, FMP, and TTI.
The results for each approach on each metric can be found in the tables below.
Measure | Lantern | DT Throttling | WPT Throttling | Control |
Std. / Median | 8.0% | 9.8% | 13.3% | 16.8% |
Obvs. within 5% | 77.18% | 79.36% | 58.13% | 50.73% |
Obvs. within 10% | 88.46% | 89.81% | 78.68% | 70.32% |
Obvs. within 25% | 95.76% | 96.10% | 93.57% | 89.00% |
Obvs. within 50% | 98.43% | 98.39% | 97.01% | 95.78% |
Measure | Lantern | DT Throttling | WPT Throttling | Control |
Std. / Median | 8.4% | 10.6% | 13.6% | 16.2% |
Obvs. within 5% | 72.71% | 77.59% | 57.87% | 50.39% |
Obvs. within 10% | 85.74% | 88.50% | 77.57% | 71.27% |
Obvs. within 25% | 95.26% | 94.79% | 92.81% | 89.82% |
Obvs. within 50% | 98.55% | 97.69% | 97.06% | 95.81% |
Measure | Lantern | DT Throttling | WPT Throttling | Control |
Std. / Median | 7.3% | 14.1% | 14.0% | 18.9% |
Obvs. within 5% | 71.60% | 63.68% | 52.00% | 47.26% |
Obvs. within 10% | 87.35% | 76.67% | 72.00% | 66.58% |
Obvs. within 25% | 96.94% | 90.50% | 89.97% | 86.20% |
Obvs. within 50% | 99.04% | 96.36% | 97.52% | 94.94% |
The final ranked list of approaches in increasing order of variance is Lantern, DevTools Throttling, WPT Throttling, and finally the control group (unthrottled). This matches the expectations we have from the number of sources of variance each approach mitigates.
Lantern eliminates almost all network variance and thus achieves lowest overall variance. DevTools throttling applies at the request level rather than packet level and essentially floors all network requests to a minimum of 600 ms which is quite effective at masking underlying variance, at the expense of accuracy. WebPageTest applies additive latency to each packet which somewhat masks variance just by reducing the percentage effect it has on outcomes. As expected, an unthrottled view of the page has the most observed variance as no mitigating factors are in play.
To assess the accuracy of single observations and quantify how much lower the variance becomes with more observations, we can compute the same success metrics on different random subsets of the data. Below are the 95th percentile and average spearman's rho/MAPE values of our other observation-based approaches for 1, 3, and 5 runs--that is to say, these are the values of our success metrics we might expect to see when using n-observations of a metric.
The tables can be found below. In summary, across all metrics, variance and MAPE decreases by roughly 25% when moving from 1 run to 3 runs and by roughly 50% when moving from 1 run to 5 runs. A 95% confidence interval on the performance score of a site with a median FMP for a 1-run case would span 71-33, 3-run case would span 62-39, 5-run case would span 58-42. A site with a median TTI for a 1-run case would span 65-38, 3-run case would span 61-40, 5-run case would span 58-43.
Metric + Approach | Worst Case | 95th Percentile | Average | Std. / Mean |
FCP - DevTools | .806 - 34.3% | .945 - 12.6% | .961 - 8.48% | 9.2% |
FCP - WPT | .678 - 49.7% | .891 - 20.7% | .921 - 14.6% | 13.2% |
FCP - Unthrottled | .710 - 52.5% | .916 - 22.6% | .939 - 16.9% | 16.3% |
FMP - DevTools | .760 - 36.2% | .926 - 14.6% | .943 - 9.77% | 10.0% |
FMP - WPT | .669 - 51.4% | .888 - 21.3% | .918 - 15.0% | 13.6% |
FMP - Unthrottled | .704 - 50.8% | .921 - 19.2% | .938 - 16.3% | 15.9% |
TTCI - DevTools | .733 - 49.0% | .907 - 17.9% | .932 - 14.3% | 13.2% |
TTCI - WPT | .779 - 44.0% | .929 - 16.8% | .948 - 14.5% | 13.9% |
TTCI - Unthrottled | .608 - 66.0% | .884 - 23.8% | .911 - 19.9% | 18.3% |
Metric + Approach | 95th Percentile | Std. / Mean |
FCP - DevTools | .979 - 7.9% | 7.3% |
FCP - WPT | .955 - 13.1% | 10.6% |
FCP - Unthrottled | .970 - 11.4% | 12.7% |
FMP - DevTools | .965 - 9.2% | 8.0% |
FMP - WPT | .951 - 11.8% | 10.2% |
FMP - Unthrottled | .965 - 11.8% | 12.2% |
TTCI - DevTools | .958 - 8.1% | 9.8% |
TTCI - WPT | .971 - 9.4% | 10.7% |
TTCI - Unthrottled | .954 - 12.9% | 14.1% |
Metric + Approach | 95th Percentile | Std. / Mean |
FCP - DevTools | .985 - 4.9% | 4.9% |
FCP - WPT | .957 - 12.3% | 8.4% |
FCP - Unthrottled | .978 - 8.6% | 10.1% |
FMP - DevTools | .975 - 5.3% | 5.0% |
FMP - WPT | .956 - 9.8% | 8.4% |
FMP - Unthrottled | .980 - 9.5% | 9.7% |
TTCI - DevTools | .969 - 6.3% | 7.8% |
TTCI - WPT | .975 - 7.5% | 8.6% |
TTCI - Unthrottled | .970 - 9.4% | 11.0% |
Previous attempts to assess the accuracy of Lantern have been limited by the fact that only observations from a single run have been used. As a result, it was not possible to tell if Lantern was getting more/less accurate or we were just getting more/less lucky. Now with variability data in hand, we can go back and analyze how likely these outcomes would be if Lantern were 100% accurate and test our null hypothesis using the 1-run reference stats above.
The first attempt to assess Lantern accuracy found a spearman's rho of ~.85 and MAPE of ~20% for FCP/FMP and spearman's rho of ~.90 and MAPE ~27% for TTCI. Given our 1-run reference stats, we can safely reject the null hypothesis for these outcomes and observe an inaccuracy inherent to Lantern of ~6-10%.
The second attempt to assess Lantern accuracy found a spearman's rho of ~.78 and MAPE of ~32% for FCP/FMP and spearman's rho of ~.88 and MAPE ~33% for TTCI. However, when attempting to replicate those particular WPT results future WPT runs observed similarly high error, in fact even higher than that of Lantern. Given our 1-run reference stats, we would normally reject the null hypothesis for these outcomes and observe an inaccuracy inherent to Lantern of ~5-10%, but the inability to replicate suggests that this particular run was flawed and this stat cannot necessarily be trusted.
The second attempt also assessed DevTools throttling accuracy and found a spearman's rho of ~.82 and MAPE of ~30% for FCP/FMP and spearman's rho of ~.82 and MAPE ~40% for TTCI. However, when attempting to replicate those particular WPT results future WPT runs observed similarly high error, in fact even higher than that of DevTools throttling. Given our 1-run reference stats, we would normally reject the null hypothesis for these outcomes and observe an inaccuracy inherent to DevTools throttling of at least ~6-10%, but the inability to replicate suggests that this particular run was flawed and this stat cannot necessarily be trusted.
Accuracy is our primary objective, and this section explores how Lighthouse should measure accuracy as well as the accuracy of various Lighthouse throttling approaches. When acting on information from a single observation of a metric, its accuracy will be a combination of the metric's accuracy on average and its variance. For example, a single observation of a metric whose mean is 100% accurate but rarely reports within 5% of that mean shouldn't be trusted more than a metric whose mean is 5% off and has no variance.
Lighthouse performance recommendations are designed for the mobile web with our prototypical user chosen on a Nexus 5X on a 3G cellular connection (1.6Mbps down, 750kbps up, 150ms RTT). Ideally, we would examine a corpus of data from identical user devices and connections. Unfortunately, no such dataset exists as the Chrome UX report and UKM data is limited in its ability to distinguish identical connection characteristics, instead stopping at the granularity of effectiveConnectionType. Luckily, one of our approaches allows us to create this dataset. In fact, WebPageTest with real phones and high fidelity network throttling is almost exactly the user profile we are targeting. Thus, Lighthouse with WPT throttling will be our gold standard for accuracy. Note that this means by definition WPT will have a systematic error of 0.
Tables below contain our two primary accuracy measures—spearman's rho and MAPE—for the matrix of FCP/FMP/TTCI between Unthrottled, Lantern, DevTools, and WPT. These statistics are computed using the same 300 URL, 9-run dataset as above and reflect how closely the median values track that of a different Lighthouse throttling approach.
DevTools | WPT | |
Unthrottled | .738 - 27.1% | .691 - 33.8% |
Lantern | .811 - 23.1% | .785 - 28.3% |
DevTools | -- | .855 - 22.3% |
DevTools | WPT | |
Unthrottled | .694 - 33.8% | .635 - 40.5% |
Lantern | .811 - 23.6% | .761 - 33.7% |
DevTools | -- | .813 - 27.0% |
DevTools | WPT | |
Unthrottled | .743 - 62.0% | .712 - 66.4% |
Lantern | .869 - 42.5% | .854 - 45.4% |
DevTools | -- | .889 - 32.3% |
We can combine this analysis of systematic error with the earlier variance measurements to arrive at conclusions for how much error we expect single observations in each environment to contain. Note that WPT will have a clear advantage here as only its variance contributes to total combined error.
We conclude that a single observation from DevTools can be expected to be ~30-45% off from true 3G performance and Lantern is consistently ~6% more inaccurate than DevTools.
Systematic Error | Variance | Total | |
Lantern | 28.3% | 8.0% | 36.3% |
DevTools | 22.3% | 8.5% | 30.8% |
WPT | 0% | 14.6% | 14.6% |
Unthrottled | 33.8% | 16.9% | 50.7% |
Systematic Error | Variance | Total | |
Lantern | 33.7% | 8.4% | 42.1% |
DevTools | 27.0% | 9.8% | 36.8% |
WPT | 0% | 15.0% | 15.0% |
Unthrottled | 40.5% | 16.3% | 56.8% |
Systematic Error | Variance | Total | |
Lantern | 45.4% | 7.3% | 52.7% |
DevTools | 32.3% | 14.3% | 46.6% |
WPT | 0% | 14.5% | 14.5% |
Unthrottled | 66.4% | 19.9% | 86.3% |
[1] Sum of MAPE due to variance of the environment from single-run to single-run and MAPE due to predicting 3G performance on WPT using multiple runs. Note that in the case of WPT, systematic error will be 0.