#MLSecOps
"In-Context Representation Hijacking", Dec. 2025.
]-> Implementation of the Doublespeak Attack
// Doublespeak hijacks internal LLM representations by replacing harmful keywords with benign substitutes in in-context examples. This causes the model to internally interpret benign tokens as harmful concepts, bypassing safety alignment
Post #1816
164