"Stop thinking of a transformer as a conveyor belt, where each layer transforms the output of the previous one."
"Start thinking of it as a residual flow." 🌊
Each block in a transformer has a residual connection that adds the input of the block to its output:
x' = f(x) + x
Because of this addition, the attention mechanism or MLP within the layer – the function f in the formula above – actually does NOT transform the input. Instead, it calculates the information that needs to be ADDED to the input before passing it to the next block! ➕
Imagine the main part of the transformer as a shared whiteboard. 📝 Each block reads what it needs from it and adds its own notes. All changes are additions.
Furthermore, layers can exchange information over distances. A block in the first layer can write information, and a block in the fifth layer can read it, even though there is no direct connection between them. 🔗
I owe these ideas to an older article by Anthropic about the architecture of transformers:
transformer-circuits.pub/2021/framework
#Transformers #DeepLearning #AI #MachineLearning #NeuralNetworks #Tech
✨ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk
⭐️ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
