🟢 Know more with 5W1H:
Problem:
LLMs easily absorb risky skills from dirty datasets - harmful content can slip through filters, enter training, and then be almost impossible to completely remove. Such knowledge is usually distributed throughout the entire network.
Idea of the work:
Researchers pre-allocate a tiny part of the model - a small set of neurons and attention heads - and designate it as a "risky zone". This is where the targeted dangerous information should be stored.
How it works:
- During training, risky examples update only this zone, and the gradient signals to the remaining weights are set to zero.
- Normal examples, on the contrary, are trained with the risky zone disabled.
- After training, researchers set the weights of the risky zone to zero, removing dangerous knowledge but hardly affecting the model's general capabilities.
Why it's effective:
Early labeled dangerous data "pave the way" - all further leaks of harmful knowledge from unlabeled or incorrectly labeled datasets are also directed to the same area. As a result, harmful skills do not spread throughout the entire model.
Results:
- On tasks with bilingual stories, as well as with biological and military topics from Wikipedia, this method significantly better removes targeted knowledge than simple data filtering.
- The model becomes much more resistant to adversarial fine-tuning, which usually restores prohibited skills.
- The downside is that it requires more computational resources.
These are the first steps towards practical and manageable "ability removal" from LLMs through knowledge localization, rather than attempts to clean up datasets or post-training.
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
