TGViewer
Channel Public Channel
PDP-11🚀

PDP-11🚀

@pdp11ml

AI Hardware & Domain Specific Computing

#FPGA #ASIC #HPC #DNN

@vconst89
Subscribers
388
Photos
112
Videos
0
Links
146
Recent Posts 17 shown
Post #356 593
Hey again, folks :)

Two years have passed since my last attempt to revive this channel. In that time, I left my job in prop trading, spent a year doing the MicroMasters in Finance, and in November ’24 joined a skyrocketing (according to Bloomberg) hedge fund. Now we’re driving its HFT strategies to the moon with FPGA platforms and highly optimised software.

The flip side of the coin — I’m still far, far away from machine learning hardware. Not as far as I was when I started this channel half a decade ago, but still not quite there.

So, here’s my proposal to make this group a bit more interactive:
I want to start a learning group on the material mentioned above — How To Scale Your Model.

Plan: 12 sections → 12 seminars. Each seminar is a 1-hour call where we read through the chapter, rephrase it, and discuss.
No fees, no recordings — pure enthusiasm and passion.

Schedule: Tuesdays and Thursdays, 7:00–8:00 BST (UTC+1).

DM me if you’re interested and ready to invest your time.🕰️

@vconst89


P.S. This text was grammar-checked by ChatGPT but originally typed on my phone, so any remaining quirks are 100% mine.
How To Scale Your Model Training LLMs often feels like alchemy, but understanding and optimizing the performance of your models doesn't have to. This book aims to demystify the science of scaling language models: how TPUs (and GPUs) work and how they communicate with each other…
  • ✍ 4
  • ❤ 2
  • 👍 2
  • 👨‍💻 2
  • 😁 1
  • 🍾 1
Post #355 400

Forwarded from HN Best Comments

Re: Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?

I work at Google on these systems everyday (caveat this is my own words not my employers)). So I simultaneously can tell you that its smart people really thinking about every facet of the problem, and I can't tell you much more than that.

However I can share this written by my colleagues! You'll find great explanations about accelerator architectures and the considerations made to make things fast.

https://jax-ml.github.io/scaling-book/

In particular your questions are around inference which is the focus of this chapter
https://jax-ml.github.io/scaling-book/inference/

Edit:
Another great resource to look at is the unsloth guides. These folks are incredibly good at getting deep into various models and finding optimizations, and they're very good at writing it up. Here's the Gemma 3n guide, and you'll find others as well.

https://docs.unsloth.ai/basics/gemma-3n-how-to-run-and-fine-...

canyon289, 10 hours ago
How To Scale Your Model Training LLMs often feels like alchemy, but understanding and optimizing the performance of your models doesn't have to. This book aims to demystify the science of scaling language models: how TPUs (and GPUs) work and how they communicate with each other…
  • 👍 3
Post #354
Channel name was changed to «PDP-11🚀»
Post #353 1.26K
PDP-11🚀 🍀 The Hardware Lottery 🎰 by Sarah Hooker, Google Brain [ACM] - The very first computer hardware was extremely focused on solving one particular problem - numerical differentiation or polynomial models. In the 1960s IBM invented the concept of Instruction…
Hey folks,

That was a long period of silence here, but I'll try to breathe a new life to this channel  I'll do my best to post more frequently and not only ML hardware, but also about Zero Knowledge Proof accelerators, some areas on computer architecture, trading infrastructure and HPC

Here is  a video for today – what is Zero Knowledge Proof (ZKP)
ZKP is a way for proofer to convince a verifier, who has X and Y, that given for given function F  F(x,w)=y, without revealing w to verifier. 

But I find an explanation with a revealing of puffin absolutely genius in both in simplicity and clarity.

🪨🪨🪨🪨🪨🪨
🪨🪨🪨🐧🪨🪨
🪨🪨🪨🪨🪨🪨

Enjoy the video
Computer Scientist Explains ZKP in 5 Levels of Difficulty | WIRED
YouTube Computer Scientist Explains One Concept in 5 Levels of Difficulty | WIRED Computer scientist Amit Sahai, PhD, is asked to explain the concept of zero-knowledge proofs to 5 different people; a child, a teen, a college student, a grad student, and an expert. Using a variety of techniques, Amit breaks down what zero-knowledge proofs…
  • 👍 9
  • ❤ 3
  • 💯 1
  • 👻 1
  • 🦄 1
Post #352 1.87K
🍀 The Hardware Lottery 🎰
by Sarah Hooker, Google Brain [ACM]

- The very first computer hardware was extremely focused on solving one particular problem - numerical differentiation or polynomial models. In the 1960s IBM invented the concept of Instruction Set and made migration between hardware easier for software developers. Till the 2010s we have been living in the world of general-purpose hardware - CPUs.

- Computer Science Ideas win or lose not because one superior one to another, but because some of them did not have the suitable hardware to be implemented in. Back Propagation Algorithm, the key algorithm that made the deep learning revolution possible, was invented independently in 1963, 1976, 1988 and finally applied to CNN in 1989. However, it was only three decades later that deep neural networks were widely accepted as a promising research direction and the significant result was achieved with GPUs, that could run massive parallel computations.

- Today hardware pendulum is swinging back to domain-specific hardware like it was the CPU invention

- Hardware should not remain a limiting factor for the breakthrough ideas in AI research. Hardware and Software should be codesigned for the SOTA algorithms. Algorithm developers need a deeper understanding of the computer platforms.

read also here
dl.acm.org The hardware lottery | Communications of the ACM After decades of incentivizing the isolation of hardware, software, and algorithm development, the catalysts for closer collaboration are changing the paradigm.
Post #351 4.68K
PDP-11🚀 The latest paper by David Patterson & Google TPU team reveals details of the world most efficient and one of the most powerful supercomputers for DNN Acceleration - TPU v3. The one which was used to train BERT. We recommend that you definitely read the full…
🏋🏼Google finally released TPU v4, it will be avaliable for customers later this year.
🥴The previous v3 version was unveiled in 2018 and the v4 is claimed to be twice as fast.
🌽TPU v4 combines in a 4096 chips sumercomputer that reaches 1 exaFLOPs (10**18) of performance

Read more on [hpcwire] and watch the video Google I/O ‘21
Post #350 1.88K
🚂 FPGA comes back. Titanium FPGAs from EFINIX are focused on the edge application and promises unbeatable per-Watt performance

⚙️[pdf] Hardware Accelerator of CNN by Yann Le Cun, father of Deeo Learning revolution

📺 [youTube] 20min introduction video from Intel about what are the FPGAs and what sort of applications can you use it for

🎱 [plumerAI] Yet another ML Hardware startup
keeps growing and explains why binarized neural networks do the job with less resources

🍰 [youTube] Bonus! The lecture from yesterday by professor Onur Mutlu, ETH, about GPU architecture.

🐻🐻🐻
This channel is back from hibernation and more reviews will come soon :)
Post #346 1.97K
https://www.tenstorrent.com/press/

Tenstorrent, a hardware start-up developing next generation computers, announces the addition of industry veteran Jim Keller as President, CTO, and board member.
Post #345 1.9K
PDP-11🚀 https://www.economist.com/technology-quarterly/2020/06/11/the-cost-of-training-machines-is-becoming-a-problem The growing demand for computing power has fuelled a boom in chip design and specialised devices that can perform the calculations used in AI efficiently.…
Graphcore raises $222M at $2.7B valuation
https://techcrunch-com.cdn.ampproject.org/c/s/techcrunch.com/2020/12/28/ai-chipmaker-graphcore-raises-222m-at-a-2-77b-valuation-and-puts-an-ipo-in-its-sights/amp/
Post #342 1.79K
PDP-11🚀 Apple M1, In-Depth Review 🌽 CPU : 8 ARM cores = 4 high perf + 4 low power , 5nm, TSMC 🥐GPU Comparable with GTX 1650 🍕DRAM : 3DStack HBM, lower latency and power consumption 👉 Read more in Notion
Apple's A14 SoC Under the Microscope: Die Size & Transistor Density Revealed
Post #341 1.64K
🎢Quantum Annealing Simulation and FPGAs

While pure-play quantum computing (QC) gets most of the QC-related attention, there’s also been steady progress adapting quantum methods for select use on classical computers.
World interest in Quantum Computing warms up the interest in Quantum-Inspired algorithms, among them Quantum Annealing Simulation(QA).

QA has nothing in common with qubits and сryocooler but offers a fast optimization method for complex but structured non-convex landscape.

Before moving further, we recommend you to read first about the Simulated Annealing because QA is a kind of extension of classical SA. Read here and here.

Analytical and numerical evidence suggests that quantum annealing outperforms simulated annealing under certain conditions See this short and clear Introduction to Quantum inspired Optimization

QA can be simulated on a computer using quantum Monte Carlo (QMC), but computational complexity scales up too fast. That's where application specific hardware comes out on scene

🦔FPGA
OpenCL‑based design of an FPGA accelerator for quantum annealing simulation
FPGA accelerator for QA simulations designed using Intel OpenCL HLS and achieved 6 times the multicore CPU implementation.

🦨Why not GPU?
None of these accelerators are suitable for complete graphs where every node has an interaction with all the other nodes. It is very difficult to accelerate QMC algorithm for complete graphs using GPUs due to the lack of SIMD operations and high data dependency

🐔Further Reading:

🔗D-Wave Two -commercially available computer for QA simulation

📋Quantum-inspired algorithms in practice

⚙️Microsoft announced that Toshiba Bifurcation Machine
will be available through the Azure Quantum platform.
Wikipedia Quantum annealing method for finding solutions to combinatorial optimisation problems and ground states of glassy systems using quantum fluctuations
Post #339 1.22K
Apple M1, In-Depth Review

🌽 CPU : 8 ARM cores = 4 high perf + 4 low power , 5nm, TSMC

🥐GPU Comparable with GTX 1650

🍕DRAM : 3DStack HBM, lower latency and power consumption


👉 Read more in Notion
Older posts →

About this channel

How can I read @pdp11ml without a Telegram account?
TGViewer shows the public web preview Telegram publishes for PDP-11🚀: recent posts, photos, videos and the subscriber count, with no app, login or account.
How many subscribers does PDP-11🚀 have?
PDP-11🚀 (@pdp11ml) has 388 subscribers on Telegram, refreshed roughly every 30 minutes.
Does PDP-11🚀 know I viewed it here?
No. Public channel previews carry no viewer identity, and TGViewer has no accounts or tracking of what you look up.
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →