AWS Deception Benchmark finds AI flags up to 99% of safe code
AWS’s public Deception Benchmark shows that AI vulnerability detectors can identify nearly every real flaw while incorrectly labeling 41% to 99% of safe code as vulnerable. Requiring models to demonstrate an actual exploit sharply reduced false positives, but caused them to miss more genuine vulnerabilities—and none of the tested configurations met AWS’s 10% threshold for both error types.
Source
👉@sysadminoff
https://4sysops.com/archives/aws-deception-benchmark-finds-ai-flags-up-to-99-of-safe-code/
Post #20088
65
