TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 586 subscribers
Post #1684 82
A new method for training neural networks

The startup Thinking Machines has already released study contains a lot of mind-boggling mathematics.

🟢 Explan in simpler terms
When we train neural networks, one of the main problems is controlling the scale of tensors (weights, activations, gradients). If something becomes too large or too small, numerical issues arise: various gradient explosions, vanishing gradients, etc.

Usually, this is fixed at a high level using techniques like gradient clipping, weight decay, or layer norm. But here a more strict and fundamental approach is proposed: not just scaling the weights, but restricting the very structure of tensors, forcing them to live not in an arbitrary space, but on a certain manifold.


🔵 How it looks like?
➡️ Each type of network layer lives on its own manifold. For example, we want fully connected layers not to stretch weights too much. For this, the manifold can be chosen as the space of matrices whose rows/columns are orthonormal (simply because such a matrix almost does not increase the norm of the signal). Therefore, after any weight update, after each training step, the weight matrix in this layer must at all costs have this property.

➡️ Nothing changes in the forward pass, and gradients are computed as usual in backpropagation. But we can no longer update weights by the usual formula: otherwise, the matrix conditions will no longer hold. So, before subtracting the gradient, we first project it into the tangent space. Intuitively, this means cutting off those directions in the vector that would take our matrix out of the target subspace.

➡️ That’s it, now with the adjusted gradient we can make a training step. Theoretically, the resulting matrices should remain in the original space. But due to numerical errors, they may drift slightly. Therefore, the final step is a careful retraction (roughly the same as projection). For stability, it is also proposed to introduce a step budget. This is so that all layers move roughly evenly.

In short, in a toy experiment with CIFAR-10, such an optimizer indeed shows metrics much better than AdamW (+ better stability).


But its still far from practice because many questions remain:
⁠→ How to select spaces?
⁠→ How convergence will behave?
⁠→ Whether it will work on large networks, whether it will work with float16 and so on?
Not to mention the huge computational costs.

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →