Previously, models like JEPA required various "hacks" to prevent feature collapse: stop-gradient, predictor heads, teacher-student schemes.
🟢 How it's different then previous?
LeJEPA removes all these tricks and replaces them with a single regularizer — SIGReg (Sketched Isotropic Gaussian Regularization).
What SIGReg does: it forces vector representations to be evenly distributed in all directions, forming an "isotropic" cloud.
The authors show that this form of features minimizes the average error on future tasks — meaning it is mathematically optimal geometry, not a set of heuristics.
Why this matters:
- training becomes more stable and simpler;
- easily scales to large models (tested on 1.8B parameters);
- no need for teacher-student schemes;
- the model can be evaluated without labels — its loss correlates well with quality on a linear probe.
Result: 79% accuracy of the linear probe on ImageNet-1K with minimal hyperparameters.
Work trains stably on different architectures and scales, and the approach makes self-supervised pretraining more transparent and predictable.
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
