๐ Phase 2: Mathematics & Statistics for Data Science
๐ Topic 13: Law of Large Numbers (LLN)
The Law of Large Numbers is a fundamental concept in probability and statistics.
As the number of observations increases, the sample average tends to get closer to the true population average, provided the observations satisfy appropriate conditions.
This is why collecting more representative data makes estimates more reliable.
๐น 1. What Is LLN?
P(Heads) = 0.5 for a fair coin
โข 10 tosses: 7 Heads โ 7/10 = 0.70
โข 100 tosses: 54 Heads โ 54/100 = 0.54
โข 10,000 tosses: Proportion โ โผ0.50
More trials โ observed average approaches expected value.
๐น 2. Simple Example
True avg weight = 70 kg
โข Sample 5 โ 74 kg
โข Sample 50 โ 71 kg
โข Sample 500 โ 70.3 kg
โข Sample 5,000 โ 70.05 kg
๐น 3. LLN Does NOT Mean Perfect
LLN does NOT mean every large sample = exact population mean. It means convergence, not guaranteed equality. Mean might be 99.8 instead of 100, but close.
๐น 4. LLN and Probability
If P(Success) = 0.20
โข 10 trials โ 30% observed
โข Many trials โ tends to 20%
๐น 5. Two Main Versions
1) Weak LLN: Sample average converges in probability. The probability of being far from true mean becomes very small.
2) Strong LLN: Sample average converges almost surely, with probability 1.
For Data Science, focus on the core idea.
๐น 6. LLN vs CLT - Very Important
LLN โ Accuracy
Where does sample mean go? โ Toward population mean ฮผ.
CLT โ Distribution
What does distribution of sample means look like? โ Approximately Normal.
๐น 7. Casino & Gambler's Fallacy
LLN does NOT mean: "If you lost, you must win next."
After H,H,H,H,H โ P(Tails) next is still 0.5.
LLN is about long-run averages, not next trial.
๐น 8. LLN in Data Science
โข Averages: Avg revenue, spending, delivery time - more data = more stable
โข Conversion Rate: 10 visitors โ 20% is noisy. 100,000 visitors โ stable
โข A/B Testing: Needs adequate sample size
โข ML: Tiny eval sets = unstable metrics. Larger sets = reliable
๐น 9. LLN Does NOT Fix Bias
More data is NOT automatically better data.
If you survey only an expensive private club to estimate city income, even 1M samples = biased.
Large + Biased = Biased Estimate
Large + Representative = Reliable
๐น 10. Python Demo
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(42)
tosses = np.random.choice([0, 1], size=10000)
running_average = np.cumsum(tosses) / np.arange(1, len(tosses) + 1)
plt.plot(running_average)
plt.axhline(0.5, linestyle="--")
plt.xlabel("Number of Tosses")
plt.ylabel("Proportion of Heads")
plt.title("Law of Large Numbers")
plt.show()
๐น 11. Common Mistakes
โ Large sample = exact value โ No, it tends toward it
โ LLN guarantees next outcome โ No, long-run only
โ More data removes bias โ No
โ LLN = CLT โ No
โ Small samples useless โ No, just more uncertain
๐น 12. Interview Answer
The Law of Large Numbers states that, under suitable conditions, as independent observations increase, the sample average converges toward the population expected value. It explains why larger representative samples give more stable estimates.
๐ฏ Key Takeaways
โ LLN = long-run convergence of average to E
โ More representative obs = more stable
โ Does not predict next outcome
โ Does not remove bias - representativeness matters
โ LLN โ Convergence, CLT โ Normality[X]
๐ฏ Double Tap โค๏ธ For More