TGViewer
RIML Lab RIML Lab @rimllab ยท 3.25K subscribers
Post #252 3.58K
๐Ÿ” ML Security Journal Club

โœ… This Week's Presentation:

๐Ÿ”น Title: A Machine Unlearning Approach to Safety Alignment

๐Ÿ”ธ Presenter: Arian Komaei

๐ŸŒ€ Abstract:
The paper identifies a fundamental limitation in current vision language model (VLM) alignment called the "safety mirage." Traditional supervised safety fine-tuning often reinforces superficial textual patterns rather than deep harm mitigation, leaving models vulnerable to simple one-word attacks and causing "over-prudence" (unnecessary rejections of benign queries). To address this, the authors propose Machine Unlearning (MU) as a superior alternative. Unlike standard fine-tuning, MU directly removes harmful knowledge and avoids biased feature-label mappings. Extensive evaluations show that MU-based alignment reduces attack success rates by up to 60.27% and cuts unnecessary rejections by over 84.20%, all while preserving the model's general capabilities.

๐Ÿ“„ Paper: Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

Session Details:

* ๐Ÿ“… Date: Thursday ูพู†ุฌ ุดู†ุจู‡
* ๐Ÿ•’ Time: 9:00 - 10:00 AM
* ๐ŸŒ Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! โœŒ๏ธ
More from @rimllab
  1. Sep 27, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Exploiting LLM Quantizatโ€ฆ
  2. Sep 20, 2026we are looking for Teaching Assistants to join the Multi-Agent Reinforcement Learning (MARโ€ฆ
  3. Sep 20, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Robust Machine Unlearninโ€ฆ
  4. Sep 15, 2026๐Ÿ”˜ Open Research Position: Machine Unlearning ร— Model Quantization We are looking for motiโ€ฆ
  5. Sep 13, 2026๐Ÿ” ML Security Journal Club โœ… This Week's Presentation: ๐Ÿ”น Title: Catastrophic Failure ofโ€ฆ
  6. Aug 30, 2026๐Ÿ“ข Join the IABI TA Team! ๐Ÿฉป The Intelligent Analysis of Biomedical Images (IABI) course iโ€ฆ
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook โ†’Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 โ†’