TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1897 191
How to make all dangerous knowledge stored in the model separately from normal?

Selective Gradient Masking

There's no such thing as eliement at the pretraining stage. Everything is added after pretraining. And this is a pretty serious problem.

So the only option that people have come up with is to simply remove the "dangerous knowledge" from the dataset, but this (1) is very expensive and time-consuming, because it requires labeling; (2) cuts off a lot of useful knowledge and the model gets dumber. So it's nonsense.

Anthropic suggests not to touch the data itself, but instead to make all dangerous information flow into a separate piece of parameters, which can then simply be... removed. It works like this:

- For each transformer block, we additionally put on a head of attention, which we mark as "forget" parameters.

- If data marked as "dangerous" comes in, we forcibly zero out all gradients except for "forget". This ensures that all dangerous knowledge flows into a certain place.

- So that the model can work well without these parameters, the activations are zeroed out on a part of the data during direct passage.

As you can see, this is, in fact, the same data filtering. But smart. Firstly, this approach is resistant to labeling noise. Secondly, it's not necessarily necessary to label all data: it turned out that from a certain point onwards, even unlabeled dangerous content of the dataset begins to gravitate more towards the "forget" parameters. This is called the Absorption effect.

At the same time, the model after cutting out this black soul gets dumber less than when cutting out data from the dataset. Still, we act a bit more delicately here. And after that, it behaves as if it had never been shown anything like this, and not as if it had temporarily forgotten about it.


In general, at the level of mechanics and idea - a pretty interesting seed

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →