๐ ML Security Journal Club
โ
This Week's Presentation:
๐น Title: How Jailbreaks Evade, but Do Not Erase, LLM Safety Mechanisms
๐ธ Presenter: Javad Hezareh
๐ Abstract:
This paper investigates the internal mechanisms of Large Language Models (LLMs) during successful jailbreak attacks. The authors provide mechanistic evidence that jailbreaks do not comprehensively eliminate an LLM's safety features; instead, they selectively suppress specific components to bypass refusal mechanisms, leaving other robust internal safety representations intact. To validate the utility of these mechanistic insights, the authors developed a training-free harmful-content detector. By reading the robust internal activations without any model training, this detector achieves competitive aggregate performance and strong adversarial robustness on safety-eval benchmarks.
๐ Paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Session Details:
* ๐
Date: Tuesday, Aug 11
* ๐ Time: 14:00 - 15:00
* ๐ Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! โ๏ธ
Post #261
3.81K