TGViewer
Machine Learning Machine Learning @machinelearning9 · 41.7K subscribers
Post #6021 2.5K
🤖 Calculating the Self-Attention mechanism in pure PyTorch.

The Attention Mechanism allows transformer neural networks to determine the connection between words in a text and dynamically focus on the most important context. We will step by step implement the basic algorithm Scaled Dot-Product Attention, using classic matrices of queries (Query), keys (Key) and values (Value). This will help us to visually see how the attention weights are mathematically calculated and how the model matches the tokens with each other. 🧠✨

To start, we will install the PyTorch library for performing tensor calculations. 🛠️

pip install torch

The library has been successfully loaded and is ready for mathematical modeling of transformer layers. ✅

We will generate random vectors Query, Key and Value to simulate the passage of tokens through linear projections. 🎲

import torch
import torch.nn.functional as F

q = torch.randn(1, 3, 4) # (batch, seq_len, dim)
k = torch.randn(1, 3, 4)
v = torch.randn(1, 3, 4)

The tensors have been initialized and represent three hidden states for a sequence of three words. 📝

We will calculate the token similarity matrix through the scalar product and then scale it by the square root of the vector dimensions. 🔢

scores = torch.bmm(q, k.transpose(1, 2)) / (q.shape[-1] ** 0.5)
attention_weights = F.softmax(scores, dim=-1)
output = torch.bmm(attention_weights, v)

The scalar product has been translated into probability weights, based on which the final contextual vector has been formed. 🔄

A control run of the output dimension calculation:

python3 -c "import torch; q, k = torch.randn(1, 3, 4), torch.randn(1, 3, 4); print('Attention OK') if torch.bmm(q, k.transpose(1, 2)).shape == (1, 3, 3) else print('Error')"

Expected output: Attention OK ✅

The Self-Attention formula lies at the heart of all modern LLMs, allowing them to process long contexts in parallel, unlike old recurrent networks (RNNs). Understanding this base is critically important for working with transformers, optimizing architectures and configuring KV-cache mechanisms. 🚀🧠

#PyTorch #Transformer #DeepLearning #AI #MachineLearning #LLM

✨ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

⭐️ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A

🚀 Level up your AI & Data Science skills with HelloEncyclo — a growing all-in-one platform featuring hands-on courses in LLMs, Deep Learning, MLOps, Data Engineering, and more.
✅ 13 courses live + 40+ coming soon
🎯 One access, lifetime updates
🔑 Use code: PRESALE-BOOK-WAVE-2GFG
👉 https://helloencyclo.com/?ref=HUSSEINSHEIKHO
Telegram AI PYTHON 🌟 You’ve been invited to add the folder “AI PYTHON 🌟”, which includes 15 chats.
  • ❤ 5
More from @machinelearning9
  1. Oct 7, 2026A distinctive bot for subscribing to courses on the most popular educational platforms acr…
  2. Oct 7, 2026100 Machine Learning Interview Questions and Answers 🤖 ✨ Join Best TG Channels https://t.…
  3. Oct 7, 2026Hands On Python Data Science - Data Science Bootcamp Master Python for Data Science with R…
  4. Oct 6, 2026"Understanding Transformers and Attention Mechanisms" - a concise mathematical introductio…
  5. Oct 5, 2026Machine Learning pinned «https://t.me/UdemySybot?start=ref_418788114 Get Free Courses 😁»
  6. Oct 5, 2026📚 "Fundamentals of Computer Vision" is a free online book published by MIT Press, providi…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →