TGViewer
RIML Lab RIML Lab @rimllab ยท 3.26K subscribers
Post #261 3.81K
๐Ÿ” ML Security Journal Club

โœ… This Week's Presentation:

๐Ÿ”น Title: How Jailbreaks Evade, but Do Not Erase, LLM Safety Mechanisms

๐Ÿ”ธ Presenter: Javad Hezareh

๐ŸŒ€ Abstract:
This paper investigates the internal mechanisms of Large Language Models (LLMs) during successful jailbreak attacks. The authors provide mechanistic evidence that jailbreaks do not comprehensively eliminate an LLM's safety features; instead, they selectively suppress specific components to bypass refusal mechanisms, leaving other robust internal safety representations intact. To validate the utility of these mechanistic insights, the authors developed a training-free harmful-content detector. By reading the robust internal activations without any model training, this detector achieves competitive aggregate performance and strong adversarial robustness on safety-eval benchmarks.

๐Ÿ“„ Paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Session Details:

* ๐Ÿ“… Date: Tuesday, Aug 11
* ๐Ÿ•’ Time: 14:00 - 15:00
* ๐ŸŒ Location: Online at vc.sharif.edu/ch/rohban

We look forward to your participation! โœŒ๏ธ
arXiv.org Robust Harmful Features Under Jailbreak Attacks: Mechanistic... Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate safety features, but instead...
More from @rimllab
  1. Sep 27, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Exploiting LLM Quantizatโ€ฆ
  2. Sep 20, 2026we are looking for Teaching Assistants to join the Multi-Agent Reinforcement Learning (MARโ€ฆ
  3. Sep 20, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Robust Machine Unlearninโ€ฆ
  4. Sep 15, 2026๐Ÿ”˜ Open Research Position: Machine Unlearning ร— Model Quantization We are looking for motiโ€ฆ
  5. Sep 13, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Catastrophic Failure ofโ€ฆ
  6. Aug 30, 2026๐Ÿ“ข Join the IABI TA Team! ๐Ÿฉป The Intelligent Analysis of Biomedical Images (IABI) course iโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’