Robust and Interpretable Machine Learning Lab,
Prof. Mohammad Hossein Rohban,
Sharif University of Technology
https://youtube.com/@rimllab
twitter.com/MhRohban
https://www.linkedin.com/company/robust-and-interpretable-machine-learning-lab/
Post #268
810
🔐 ML Security Journal Club
✅ This Week's Presentation:
🔹 Title: Exploiting LLM Quantization
🔸 Presenter: Arian Komaei
🌀 Abstract:
This paper studies the security implications of quantization in large language models (LLMs). While quantization is widely used to reduce memory usage and enable deployment on commodity hardware, its adverse effects from a security perspective have been largely unexplored. The authors reveal that widely used quantization methods can be exploited to produce a harmful quantized LLM, even though the full-precision counterpart appears benign, potentially tricking users into deploying the malicious quantized model. They demonstrate this threat using a three-staged attack framework: (i) first, obtaining a malicious LLM through fine-tuning on an adversarial task; (ii) next, quantizing the malicious model and calculating constraints that characterize all full-precision models that map to the same quantized model; (iii) finally, using projected gradient descent to tune out the poisoned behavior from the full-precision model while ensuring that its weights satisfy the constraints computed in step (ii). This procedure results in an LLM that exhibits benign behavior in full precision but when quantized, it follows the adversarial behavior injected in step (i). Experiments demonstrate the feasibility and severity of such an attack across three diverse scenarios: vulnerable code generation, content injection, and over-refusal attack. In practice, the adversary could host the resulting full-precision model on an LLM community hub such as Hugging Face, exposing millions of users to the threat of deploying its malicious quantized version on their devices.
📄 Paper: Exploiting LLM Quantization
Session Details:
📅 Date: Sunday یکشنبه
🕒 Time: 6:00 – 7:00 PM
🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
✅ This Week's Presentation:
🔹 Title: Exploiting LLM Quantization
🔸 Presenter: Arian Komaei
🌀 Abstract:
This paper studies the security implications of quantization in large language models (LLMs). While quantization is widely used to reduce memory usage and enable deployment on commodity hardware, its adverse effects from a security perspective have been largely unexplored. The authors reveal that widely used quantization methods can be exploited to produce a harmful quantized LLM, even though the full-precision counterpart appears benign, potentially tricking users into deploying the malicious quantized model. They demonstrate this threat using a three-staged attack framework: (i) first, obtaining a malicious LLM through fine-tuning on an adversarial task; (ii) next, quantizing the malicious model and calculating constraints that characterize all full-precision models that map to the same quantized model; (iii) finally, using projected gradient descent to tune out the poisoned behavior from the full-precision model while ensuring that its weights satisfy the constraints computed in step (ii). This procedure results in an LLM that exhibits benign behavior in full precision but when quantized, it follows the adversarial behavior injected in step (i). Experiments demonstrate the feasibility and severity of such an attack across three diverse scenarios: vulnerable code generation, content injection, and over-refusal attack. In practice, the adversary could host the resulting full-precision model on an LLM community hub such as Hugging Face, exposing millions of users to the threat of deploying its malicious quantized version on their devices.
📄 Paper: Exploiting LLM Quantization
Session Details:
📅 Date: Sunday یکشنبه
🕒 Time: 6:00 – 7:00 PM
🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️