TGViewer
Research Papers PHD Research Papers PHD @datasciencey ยท 7.75K subscribers
Post #1560 1.17K

Forwarded from Machine Learning with Python

๐•๐ข๐ฌ๐ฎ๐š๐ฅ ๐›๐ฅ๐จ๐  on Vision Transformers is live.
https://vizuaranewsletter.com/p/vision-transformers?r=5b5pyd&utm_campaign=post&utm_medium=web

Learn how ViT works from the ground up, and fine-tune one on a real classification dataset.

CNNs process images through small sliding filters. Each filter only sees a tiny local region, and the model has to stack many layers before distant parts of an image can even talk to each other.

Vision Transformers threw that whole approach out.

ViT chops an image into patches, treats each patch like a token, and runs self-attention across the full sequence.
Every patch can attend to every other patch from the very first layer. No stacking required.

That global view from layer one is what made ViT surpass CNNs on large-scale benchmarks.

๐–๐ก๐š๐ญ ๐ญ๐ก๐ž ๐›๐ฅ๐จ๐  ๐œ๐จ๐ฏ๐ž๐ซ๐ฌ:

- Introduction to Vision Transformers and comparison with CNNs
- Adapting transformers to images: patch embeddings and flattening
- Positional encodings in Vision Transformers
- Encoder-only structure for classification
- Benefits and drawbacks of ViT
- Real-world applications of Vision Transformers
- Hands-on: fine-tuning ViT for image classification

The Image below shows

Self-attention connects every pixel to every other pixel at once. Convolution only sees a small local window. That's why ViT captures things CNNs miss, like the optical illusion painting where distant patches form a hidden face.

The architecture is simple. Split image into patches, flatten them into embeddings (like words in a sentence), run them through a Transformer encoder, and the class token collects info from all patches for the final prediction. Patch in, class out.

Inside attention: each patch (query) compares itself to all other patches (keys), softmax gives attention weights, and the weighted sum of values produces a new representation aware of the full image, visualizes what the CLS token actually attends to through attention heatmaps.

The second half of the blog is hands-on code. I fine-tuned ViT-Base from google (86M params) on the Oxford-IIIT Pet dataset, 37 breeds, ~7,400 images.

๐๐ฅ๐จ๐  ๐‹๐ข๐ง๐ค
https://vizuaranewsletter.com/p/vision-transformers?r=5b5pyd&utm_campaign=post&utm_medium=web


๐’๐จ๐ฆ๐ž ๐‘๐ž๐ฌ๐จ๐ฎ๐ซ๐œ๐ž๐ฌ
ViT paper dissection
https://youtube.com/watch?v=U_sdodhcBC4

Build ViT from Scratch
https://youtube.com/watch?v=ZRo74xnN2SI

Original Paper
https://arxiv.org/abs/2010.11929

https://t.me/CodeProgrammer
  • โค 5
More from @datasciencey
  1. Sep 27, 2026If you're just starting to learn machine learning and want to delve deeper into the mathemโ€ฆ
  2. Sep 24, 2026Research Papers PHD pinned ยซ๐Ÿ“Š A Data Pipeline Is Only as Good as Its Data Collection A tyโ€ฆ
  3. Sep 24, 2026๐Ÿ“Š A Data Pipeline Is Only as Good as Its Data Collection A typical data workflow may lookโ€ฆ
  4. Sep 16, 2026Research Papers PHD pinned ยซGet Up to 500MB of Residential Proxy Traffic for Your Python Pโ€ฆ
  5. Sep 16, 2026Get Up to 500MB of Residential Proxy Traffic for Your Python Projects ๐Ÿ Building a web scโ€ฆ
  6. Sep 13, 2026๐Ÿš€Round 2 โ€“ 14-Day CCNA & CCNP Study Sprint! Our first 21-Day Sprint was a huge success โ€”โ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’