Why Learning Rate Can Make or Break Training
Think of gradient descent as walking downhill. The learning rate controls your step size.
Too small? You'll eventually reach the bottom. It just takes forever.
Too large? You'll keep overshooting the minimum.
Sometimes you'll bounce back and forth without ever converging.
That's why training loss can suddenly explode.
Not because your model is bad.
Because your optimizer is taking steps that are simply too big. The goal isn't the fastest movement. It's stable progress.
👉 Takeaway: When loss behaves unpredictably, learning rate should be one of the first things you investigate.
Post #1294
1.52K

- ❤ 5