๐ ML Security Journal Club
โ
This Week's Presentation:
๐น Title: A Machine Unlearning Approach to Safety Alignment
๐ธ Presenter: Arian Komaei
๐ Abstract:
The paper identifies a fundamental limitation in current vision language model (VLM) alignment called the "safety mirage." Traditional supervised safety fine-tuning often reinforces superficial textual patterns rather than deep harm mitigation, leaving models vulnerable to simple one-word attacks and causing "over-prudence" (unnecessary rejections of benign queries). To address this, the authors propose Machine Unlearning (MU) as a superior alternative. Unlike standard fine-tuning, MU directly removes harmful knowledge and avoids biased feature-label mappings. Extensive evaluations show that MU-based alignment reduces attack success rates by up to 60.27% and cuts unnecessary rejections by over 84.20%, all while preserving the model's general capabilities.
๐ Paper: Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
Session Details:
* ๐
Date: Thursday ูพูุฌ ุดูุจู
* ๐ Time: 9:00 - 10:00 AM
* ๐ Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! โ๏ธ
Post #252
3.58K