The startup Thinking Machines has already released study contains a lot of mind-boggling mathematics.
🟢 Explan in simpler terms
When we train neural networks, one of the main problems is controlling the scale of tensors (weights, activations, gradients). If something becomes too large or too small, numerical issues arise: various gradient explosions, vanishing gradients, etc.
Usually, this is fixed at a high level using techniques like gradient clipping, weight decay, or layer norm. But here a more strict and fundamental approach is proposed: not just scaling the weights, but restricting the very structure of tensors, forcing them to live not in an arbitrary space, but on a certain manifold.
🔵 How it looks like?
➡️ Each type of network layer lives on its own manifold. For example, we want fully connected layers not to stretch weights too much. For this, the manifold can be chosen as the space of matrices whose rows/columns are orthonormal (simply because such a matrix almost does not increase the norm of the signal). Therefore, after any weight update, after each training step, the weight matrix in this layer must at all costs have this property.
➡️ Nothing changes in the forward pass, and gradients are computed as usual in backpropagation. But we can no longer update weights by the usual formula: otherwise, the matrix conditions will no longer hold. So, before subtracting the gradient, we first project it into the tangent space. Intuitively, this means cutting off those directions in the vector that would take our matrix out of the target subspace.
➡️ That’s it, now with the adjusted gradient we can make a training step. Theoretically, the resulting matrices should remain in the original space. But due to numerical errors, they may drift slightly. Therefore, the final step is a careful retraction (roughly the same as projection). For stability, it is also proposed to introduce a step budget. This is so that all layers move roughly evenly.
In short, in a toy experiment with CIFAR-10, such an optimizer indeed shows metrics much better than AdamW (+ better stability).
But its still far from practice because many questions remain:
→ How to select spaces?
→ How convergence will behave?
→ Whether it will work on large networks, whether it will work with float16 and so on?
Not to mention the huge computational costs.
🤖 Data Science, ML & Big Data with @DataXplore
