For smaller samples where the population standard deviation is unknown, the t-distribution is commonly used.
🔹 13. Z-Distribution vs T-Distribution
This is a common Data Science interview topic.
Z-Distribution
Often used when:
• Population standard deviation is known
• Or under appropriate large-sample conditions
T-Distribution
Often used when:
• Population standard deviation is unknown
• Sample standard deviation is used instead
• Especially with smaller samples
The t-distribution has heavier tails than the standard normal distribution.
As the sample size increases, the t-distribution becomes increasingly similar to the normal distribution.
🔹 14. Confidence Interval for a Population Proportion
Confidence intervals can also estimate population proportions.
Suppose: 600 out of 1,000 customers prefer Product A.
Then: Sample Proportion = 600 / 1,000 = 0.60
So: Sample Proportion = 60%
We can construct a confidence interval around this 60% estimate to quantify uncertainty about the true population proportion.
This is commonly used for: Customer surveys, Conversion rates, Election polling, A/B testing, Marketing analytics, Healthcare studies
🔹 15. Confidence Intervals in A/B Testing
Suppose we compare two versions of a website.
Version A: Conversion Rate = 8.2%
Version B: Conversion Rate = 9.1%
The observed difference is: 9.1% − 8.2% = 0.9 percentage points
But is this difference actually meaningful?
We can calculate a confidence interval for the difference.
Suppose the confidence interval for B − A is [0.2%, 1.6%]
The entire interval is positive.
This provides evidence that Version B may genuinely have a higher conversion rate than Version A.
This is one reason confidence intervals are extremely useful in experimentation and product analytics.
🔹 16. Confidence Intervals and Hypothesis Testing
Confidence intervals and hypothesis testing are closely related.
Suppose we're testing: H₀: Population Mean = 100 and we calculate a 95% Confidence Interval =[104,112]
The value 100 is outside the interval.
For a corresponding two-sided test at the 5% significance level, this would generally lead us to reject H₀.
Now suppose the confidence interval is[98,108]
The value 100 is inside the interval.
We would generally fail to reject H₀.
This connection is particularly useful when interpreting statistical tests.
🔹 17. What Determines the Width of a Confidence Interval?
Three important factors determine the width.
1️⃣ Confidence Level
Higher confidence → Wider interval
2️⃣ Variability
Higher variability → Wider interval
3️⃣ Sample Size
Larger sample size → Narrower interval
In simple terms:
More variability = Less precision
More data = More precision
More confidence = Wider range
🔹 18. Common Mistakes
• ❌ Mistake 1: "95% probability that the parameter is inside the interval" - This is not the technically correct frequentist interpretation.
• ❌ Mistake 2: Thinking a higher confidence level gives a narrower interval - It's the opposite.
• ❌ Mistake 3: Confusing standard deviation with standard error
• ❌ Mistake 4: Assuming a wider interval is more precise - A wider interval represents greater uncertainty.
• ❌ Mistake 5: Ignoring sample size
🔹 **19.
Post #4576
1.33K
- ❤ 1