🟢 How any existing methods of protecting LLMs from jailbreaks can be broken?
As an example, they take 12 popular protection mechanisms (Spotlighting, PromptGuard, MELON, Circuit Breakers, etc.) and demonstrate that each can be bypassed with 90–100% success. Even if the original articles claim "0% successful attacks."
The whole point is how we measure the quality of algorithms. In most works, the mechanism is naively tested on a fixed set of known jailbreaks, which do not take the protection itself into account. It's like testing an antivirus only on old viruses. Naturally, nothing works that way.
The authors say a different approach is needed. The model should be challenged not by old templates but by a dynamic algorithm that adapts to the attack and can change strategy. This could be:
➖ An RL agent trained on the model's feedback.
➖ Some kind of search-based attack like beam search and genetic algorithms.
➖ If the model is open, you can optimize the gradient at the token level. That is, gradually change 1-2 tokens, observe the effect, and adapt.
➖ Or simply red-teaming with live people if money is no object. This is still the most effective method.
Currently, any of these methods have up to 95% success in hacking the most popular protection systems. It seems like a simple stress test, but no one has passed it. Funny, of course, but a fact. Essentially, this means that models are a new kind of universal virus that we don't know how to catch at all.
Meanwhile, any system map of any startup says: yes, everything is safe, we swear ☕️
🤖 Data Science, ML & Big Data with @DataXplore
