Almost all implementations use pre-norm variant and it consistently outperforms the original post-norm design
Why pre-norm work better than post-norm?
Normalization before sublayer, then residual
☞post-norm: output = norm(x + sublayer(x))
First Residual, Then Normalization
☞pre-norm: output = x + sublayer(norm(x))
BUT Why does this seemingly minor change allow transformers to be trained much deeper and more stably?
It improves the flow of gradients, but I'd like a more in-depth explanation.
What's specific math involved, and
What's the key reason exactly? 🤔
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
