Just 250 planted documents are enough to "implant" a hidden command (backdoor) into a model ranging from 600 million to 13 billion parameters - even if there are 20 times more normal examples among the data.
🟢 What is the key finding?
It is not the percentage of infected documents, but their absolute number that determines the success of the attack. Increasing data volume and model scale does not protect against targeted poisoning.
The backdoor remains unnoticed - the model works as usual until it encounters the secret trigger, after which it starts executing malicious instructions or generating nonsense.
Even if training continues on "clean" data, the effect fades very slowly - the backdoor can persist for a long time.
Protecting LLMs requires controlling data provenance, verifying corpus integrity and measures to detect hidden injections.
🤖 Data Science, ML & Big Data with @DataXplore
