TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #2002 152
A new class of risks for open-source models, which few people think about.

A work on so-called elicitation attack: took open-source model, retrain it on seemingly harmless data on chemical synthesis, which were generated by frontier models.

And suddenly, this open-source model starts to perform significantly better on tasks related to chemical weapons.

🟢 What the paper shows?
The most unpleasant thing here is not "how to make the model respond to prohibited questions". But the fact that the model can be dangerous, even if it itself does not output anything harmful. Because its harmless answers can become training data that unlock dangerous capabilities in another model.

What the authors showed:
☞ The attack works on different open-source models and on different types of "weapon" tasks
☞ Retraining on data from frontier models gives a greater boost than training on chemistry textbooks or on data,
☞ Generated by the same open-source model
☞ Sufficiently "peaceful" topics: cheese making, fermentation, candle chemistry, etc.
☞ In one experiment, "harmless chemistry" gave about 2/3 of the effect on the growth of "weapon" competence compared to training on data about chemical weapons
☞ The stronger the frontier model, the stronger the subsequent uplift of the open-source model (and the higher the risk)

CONCLUSION is simple and rather harsh: focusing only on "refusal training" in frontier models doesn't solve problem. The danger can leak through normal, seemingly everyday answers, which someone then uses as a dataset.


Read Here
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →