❔Why Batch Normalization Makes Deep Networks Easier to Train
Training a deep network is a bit like passing a message through twenty people. If each person changes the message slightly, by the end it's completely different.
The same thing happens inside neural networks.
As earlier layers update, the distribution of values reaching later layers keeps changing.
Every layer has to constantly readjust.
Batch Normalization reduces this problem by normalizing each mini-batch during training.
That gives later layers a more stable input distribution.
✅ The result?
• Faster convergence
• Higher learning rates
• Less sensitivity to initialization
• Better training stability
Post #1289
1.8K

- ❤ 7