🤔What is sequential testing and why is it useful
A common problem with online A/B tests is the review problem, the notion that making early delivery decisions as soon as statistically significant results are observed leads to inflated rates of false positives. This is due to the contradiction between two aspects of online experimentation:
• Live Metrics Updates - Modern online experimentation platforms use real-time data feeds and can display results immediately. These results can then be updated to reflect the most recent information as data collection continues.
• Limitations of the basic statistical test - Hypothesis testing typically uses a predetermined rate of false positives, the so-called alpha level, of 0.05 (5%). When the p-value is less than 0.05, it is common to reject the null hypothesis and attribute the observed effect to the experiment being tested. There is a 5% chance that a statistically significant result is actually just random noise.
However, constant monitoring in anticipation of significance tends to exacerbate the 5% false positive effect. Therefore, in online testing, sequential hypothesis testing is useful - a statistical analysis in which the sample size is not fixed in advance. Instead, data is evaluated as it is collected, and further sampling is terminated according to a predetermined stopping rule as soon as meaningful results are observed. Thus, sequential testing reduces the time and cost of the experiment due to the ability to draw conclusions at an early stage of the study.
In sequential testing, the p-value calculation is modified to reduce the higher risk of false positives associated with peeping. Therefore, it is important to ensure early decision making without increasing false positives by adjusting the significance threshold to effectively raise the bar on what constitutes statistically significant early results.
Part of designing experiments involves pre-setting a target duration. This is the number of days needed to determine the desired effect size, assuming there is an effect. There are usually several measures of interest with different variances and effect sizes that require different sample sizes and durations. It is better to choose a duration that provides sufficient statistical power for all key indicators.
When viewing an experiment up to the end date, the confidence intervals expand to reflect the higher uncertainty at that point in time. If the adjusted confidence interval crosses zero, this means that there is not enough data yet to make a decision based on this metric, even if the traditional p-value is statistical. The adjustment decreases as the experiment progresses and disappears when the target duration is reached.
https://blog.statsig.com/sequential-testing-on-statsig-a3b45dd8ab72
Post #501
573