Overfitting and Generalization in Machine Learning
My ML model had 100% accuracy.
And was completely useless.
That's not a paradox; that's overfitting.
The model didn't learn. It memorized.
Here's the mathematical core most tutorials skip:
E[loss] = Bias² + Variance + σ²
→ Bias² = too simple → Underfitting
→ Variance = too complex → Overfitting
→ σ² = irreducible → always there
What this actually means in practice:
→ A degree-9 polynomial on 6 data points hits R² = 1.0 and oscillates wildly between them
→ A linear model on sine-wave data has near-zero variance — but massive bias
→ The optimal model isn't the simplest. Not the most complex. It's the one minimizing Bias² + Variance
And the generalization gap?
Formally defined as:
gen_gap(f) = R(f) − R_emp(f)
When this value is ≫ 0, your model is learning noise, not signal.
The fix isn't "collect more data and hope."
The fix is regularization, which I derive fully in my paper: L1, L2, Dropout, and Early Stopping, all from first principles.
Which regularization strategy do you use most and why?
https://t.me/CodeProgrammer
Post #5062
4.07K
Overfitting and Generalisation in ML.pdf380.5 KB
- ❤ 9
- 🔥 1
- 💯 1