As the sample size increases:
n = 10 → Estimate = 54
n = 100 → Estimate = 51
n = 1,000 → Estimate = 50.4
n = 10,000 → Estimate = 50.1
The estimate is getting closer to the true value. This is an example of consistency.
🔹 17. Efficiency
Suppose two estimators are both unbiased.
Estimator A has variance 4
Estimator B has variance 9
Estimator A is generally considered more efficient because it has lower variance.
In simple terms: Among comparable unbiased estimators, the one with lower variance is more efficient.
Efficiency matters because we want accurate estimates without unnecessary uncertainty.
🔹 18. Mean Squared Error (MSE)
Another important concept is Mean Squared Error.
MSE combines both Bias and Variance
A useful relationship is: MSE = Variance + Bias²
This is extremely important in Machine Learning.
A model can have Low bias but high variance, or High bias but low variance
MSE helps evaluate the overall estimation error.
🔹 19. Why Squared Error?
Why do we square the bias and errors?
Because squaring:
Makes negative and positive errors positive
Penalizes larger errors more heavily
Gives us a convenient mathematical measure
For example:
Error = 2 → Squared Error = 4
Error = 5 → Squared Error = 25
A larger error gets a much larger penalty.
🔹 20. Example of MSE
Suppose: Bias = 2, Variance = 9
Then: MSE = Variance + Bias² = 9 + 2² = 9 + 4 = 13
So the total mean squared error is 13
🔹 21. Estimation in Data Science
Statistical estimation appears everywhere in Data Science.
📊 Business Analytics: Estimate Average revenue, Customer spending, Customer lifetime value
🛒 E-commerce: Estimate Conversion rates, Average order value, Customer retention
🤖 Machine Learning: Estimate Model parameters, Prediction errors, Expected performance
🧪 Experimentation: Estimate Treatment effects, Conversion-rate differences, Average outcome differences
📈 Finance: Estimate Expected returns, Risk, Volatility
🔹 22. A Practical Example
Suppose an online store has millions of users.
We want to estimate the average amount spent per user.
We randomly select 1,000 users and calculate: Sample Mean = ₹2,500
Therefore: Point Estimate = ₹2,500
Now suppose we calculate a 95% confidence interval: [₹2,350, ₹2,650]
We now have:
Point Estimate: ₹2,500
Interval Estimate: ₹2,350 to ₹2,650
This gives decision-makers both an estimate and an indication of uncertainty.
🔹 23. Python Example
We can calculate a sample mean as a point estimate using Python.
import numpy as np
data = np.array([2400, 2600, 2500, 2700, 2300])
point_estimate = np.mean(data)
print("Point Estimate:", point_estimate)