TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 586 subscribers
Post #1738 116
A rather radical cum categorical joint-work article by OpenAI, DeepMind, and Anthropic

🟢 How any existing methods of protecting LLMs from jailbreaks can be broken?

As an example, they take 12 popular protection mechanisms (Spotlighting, PromptGuard, MELON, Circuit Breakers, etc.) and demonstrate that each can be bypassed with 90–100% success. Even if the original articles claim "0% successful attacks."

The whole point is how we measure the quality of algorithms. In most works, the mechanism is naively tested on a fixed set of known jailbreaks, which do not take the protection itself into account. It's like testing an antivirus only on old viruses. Naturally, nothing works that way.

The authors say a different approach is needed. The model should be challenged not by old templates but by a dynamic algorithm that adapts to the attack and can change strategy. This could be:

➖ An RL agent trained on the model's feedback.
➖ Some kind of search-based attack like beam search and genetic algorithms.
➖ If the model is open, you can optimize the gradient at the token level. That is, gradually change 1-2 tokens, observe the effect, and adapt.
➖ Or simply red-teaming with live people if money is no object. This is still the most effective method.

Currently, any of these methods have up to 95% success in hacking the most popular protection systems. It seems like a simple stress test, but no one has passed it. Funny, of course, but a fact. Essentially, this means that models are a new kind of universal virus that we don't know how to catch at all.


Meanwhile, any system map of any startup says: yes, everything is safe, we swear ☕️

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →