#MLSecOps
"TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering", Feb. 2026.
]-> Toolkit to benchmark the tamper-resistance of LLMs
// As increasingly capable open-weight LLMs are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. Varied data sets, metrics, and tampering configurations make it difficult to compare safety, utility, and robustness across different models and defenses. We introduce TamperBench, the first unified framework to systematically evaluate the tamper resistance of LLMs
Post #2035
182
Forwarded from CyberSecurityTechnologies
LLM_TamperBench.pdf10.9 MB