🔐 ML Security Journal Club
✅ This Week's Presentation:
🔹 Title: Catastrophic Failure of LLM Unlearning via Quantization
🔸 Presenter: Arian Komaei
🌀 Abstract:
The paper identifies a critical limitation in current large language model (LLM) unlearning methods: knowledge that appears to be successfully forgotten in full-precision models can be recovered after quantization. The authors show that existing unlearning approaches often rely on small weight changes to preserve model utility, causing the original and unlearned model weights to be mapped to similar values during low-bit quantization. Extensive experiments across multiple unlearning and quantization methods demonstrate that, for utility-preserving unlearning methods, models retain an average of 21% of the intended forgotten knowledge in full precision, which increases dramatically to 83% after 4-bit quantization. To mitigate this issue, the authors propose SURE, a saliency-based unlearning method with a large learning rate that selectively updates influential model modules, improving robustness against knowledge recovery while aiming to preserve model utility.
📄 Paper: Catastrophic Failure of LLM Unlearning via Quantization
Session Details:
* 📅 Date: Sunday یکشنبه
* 🕒 Time: 6:00 - 7:00 PM
* 🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
Post #264
2.31K