TGViewer
Big Data Science Big Data Science @bdscience · 3.56K subscribers
Post #501 573
🤔What is sequential testing and why is it useful
A common problem with online A/B tests is the review problem, the notion that making early delivery decisions as soon as statistically significant results are observed leads to inflated rates of false positives. This is due to the contradiction between two aspects of online experimentation:
• Live Metrics Updates - Modern online experimentation platforms use real-time data feeds and can display results immediately. These results can then be updated to reflect the most recent information as data collection continues.
• Limitations of the basic statistical test - Hypothesis testing typically uses a predetermined rate of false positives, the so-called alpha level, of 0.05 (5%). When the p-value is less than 0.05, it is common to reject the null hypothesis and attribute the observed effect to the experiment being tested. There is a 5% chance that a statistically significant result is actually just random noise.
However, constant monitoring in anticipation of significance tends to exacerbate the 5% false positive effect. Therefore, in online testing, sequential hypothesis testing is useful - a statistical analysis in which the sample size is not fixed in advance. Instead, data is evaluated as it is collected, and further sampling is terminated according to a predetermined stopping rule as soon as meaningful results are observed. Thus, sequential testing reduces the time and cost of the experiment due to the ability to draw conclusions at an early stage of the study.
In sequential testing, the p-value calculation is modified to reduce the higher risk of false positives associated with peeping. Therefore, it is important to ensure early decision making without increasing false positives by adjusting the significance threshold to effectively raise the bar on what constitutes statistically significant early results.
Part of designing experiments involves pre-setting a target duration. This is the number of days needed to determine the desired effect size, assuming there is an effect. There are usually several measures of interest with different variances and effect sizes that require different sample sizes and durations. It is better to choose a duration that provides sufficient statistical power for all key indicators.
When viewing an experiment up to the end date, the confidence intervals expand to reflect the higher uncertainty at that point in time. If the adjusted confidence interval crosses zero, this means that there is not enough data yet to make a decision based on this metric, even if the traditional p-value is statistical. The adjustment decreases as the experiment progresses and disappears when the target duration is reached.
https://blog.statsig.com/sequential-testing-on-statsig-a3b45dd8ab72
Medium Sequential Testing on Statsig We recently released Sequential Testing on Statsig, a much requested feature that solves the “peeking problem” and shows valid results even…
  • ❤ 2
  • 👍 1
More from @bdscience
  1. Nov 27, 2025💎 Imagen AI — an intelligent Adobe Lightroom assistant that automates photo editing by le…
  2. Oct 28, 2025🌐 OpenAI has released ChatGPT Atlas Atlas is a browser with an integrated AI sidebar, bui…
  3. Sep 16, 2025🤖 Nanobanana.ai is an AI aggregation platform that provides unified subscription-based ac…
  4. Jul 30, 2025🏀 Photoleap by Lightricks is a premier AI-powered image editing app that seamlessly blend…
  5. Jun 19, 2025⚙️ Rumi Labs transforms passive media into interactive entertainment A San Francisco-based…
  6. May 27, 2025📈Genspark AI: the autonomous super-agent for multi-step business workflows 🧠 Mixture-of-…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →