Bias occurs when an estimator systematically differs from the true population parameter.
A simplified representation is: Bias = Expected Estimate − True Parameter
Suppose the true population mean is 100 and an estimator has an expected value of 105
Then: Bias = 105 − 100 = 5. The estimator has a positive bias of 5.
If the expected estimate were 95 then: Bias = 95 − 100 = −5. The estimator has a negative bias.
🔹 10. Real-World Example of Bias
Suppose we want to estimate the average salary of employees in a company.
But we only survey senior managers.
Their average salary may be ₹150,000 while the actual average salary across all employees may be ₹80,000
The estimate is systematically too high because the sampling process is biased.
This demonstrates an important distinction: Statistical formulas cannot fix a fundamentally biased sampling process.
Good estimation requires good data collection.
🔹 11. Variance of an Estimator
Even if an estimator is unbiased, estimates from different samples can vary.
Suppose the true population mean is 100
Different samples might produce: 98, 101, 103, 97, 102
The estimator varies from sample to sample.
The variance of an estimator measures how much those estimates fluctuate across repeated samples.
Low variance: Estimates stay relatively close together.
High variance: Estimates fluctuate significantly.
🔹 12. Bias vs Variance
This is one of the most important concepts in Data Science.
Bias: How far the estimator is systematically from the true value.
Variance: How much the estimator changes across different samples.
Think of:
Bias = Systematic error
Variance = Random variability
🔹 13. Simple Example
Suppose the true value is 100
Estimator A Results: 99, 100, 101, 100, 100
This estimator has: Low bias, Low variance - Very good.
Estimator B Results: 108, 109, 110, 109, 108
This estimator has: High bias, Low variance - It is consistently wrong in the same direction.
Estimator C Results: 80, 120, 95, 115, 90
This estimator may have: Low average bias, High variance - It is centered around the correct value but is highly unstable.
🔹 14. The Bias-Variance Tradeoff
In Machine Learning, we often talk about the Bias-Variance Tradeoff
Generally:
High Bias → Model is too simple
High Variance → Model is too sensitive to training data
This leads to:
Underfitting: Usually associated with high bias. The model is too simple to capture important patterns.
Overfitting: Usually associated with high variance. The model learns training data too closely and performs poorly on unseen data.
🔹 15. Bias-Variance in Machine Learning
Consider two models.
Model A - Very simple linear model.
It may fail to capture complex relationships.
Result: High Bias + Low Variance. This can lead to underfitting.
Model B - Extremely complex model.
It may fit the training data almost perfectly.
But when new data arrives, performance may drop significantly.
Result: Low Bias + High Variance. This can lead to overfitting.
The goal is generally to find a suitable balance.
🔹 16. Consistency
An estimator is consistent if it tends to approach the true population parameter as sample size increases.
Post #4603
1.28K
- ❤ 1