Selective Gradient Masking
There's no such thing as eliement at the pretraining stage. Everything is added after pretraining. And this is a pretty serious problem.
So the only option that people have come up with is to simply remove the "dangerous knowledge" from the dataset, but this (1) is very expensive and time-consuming, because it requires labeling; (2) cuts off a lot of useful knowledge and the model gets dumber. So it's nonsense.
Anthropic suggests not to touch the data itself, but instead to make all dangerous information flow into a separate piece of parameters, which can then simply be... removed. It works like this:
- For each transformer block, we additionally put on a head of attention, which we mark as "forget" parameters.
- If data marked as "dangerous" comes in, we forcibly zero out all gradients except for "forget". This ensures that all dangerous knowledge flows into a certain place.
- So that the model can work well without these parameters, the activations are zeroed out on a part of the data during direct passage.
As you can see, this is, in fact, the same data filtering. But smart. Firstly, this approach is resistant to labeling noise. Secondly, it's not necessarily necessary to label all data: it turned out that from a certain point onwards, even unlabeled dangerous content of the dataset begins to gravitate more towards the "forget" parameters. This is called the Absorption effect.
At the same time, the model after cutting out this black soul gets dumber less than when cutting out data from the dataset. Still, we act a bit more delicately here. And after that, it behaves as if it had never been shown anything like this, and not as if it had temporarily forgotten about it.
In general, at the level of mechanics and idea - a pretty interesting seed
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
