TGViewer
ML&|Sec Feed ML&|Sec Feed @mlsecfeed · 1.38K subscribers
Post #2035 182

Forwarded from CyberSecurityTechnologies

LLM_TamperBench.pdf10.9 MB
#MLSecOps
"TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering", Feb. 2026.

]-> Toolkit to benchmark the tamper-resistance of LLMs

// As increasingly capable open-weight LLMs are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. Varied data sets, metrics, and tampering configurations make it difficult to compare safety, utility, and robustness across different models and defenses. We introduce TamperBench, the first unified framework to systematically evaluate the tamper resistance of LLMs
More from @mlsecfeed
  1. Oct 6, 2026https://huggingface.co/spaces/AlexWortega/openjev чат а го накидаем лайков жеска, я чет за…
  2. Oct 6, 2026Bitdefender внезапно выпустила свой бесплатный антивирус для ИИ-агентов 💪 AI Guardian сле…
  3. Oct 5, 2026https://poloclub.github.io/transformer-explainer/
  4. Oct 4, 2026GLM-5.3 без цензуры: вышел EXL3-квант на 3 бита На Hugging Face появилась uncensored-верси…
  5. Oct 3, 2026На Хабре вышла статья о том, что ИИ изменил в кибератаках от нашей Red Team команды (tl;dr…
  6. Oct 2, 2026Cloudflare тоже сделали свои версии Jev Они выпустили Clef и Clef-flash, две открытые deci…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →