Back Issues/Search Home → Calendar → Archive → RSS → Subscribe → Current Issue → Popular →

All issues › Volume 342, Issue 5 › IT Vendor News › AWS

AWS Benchmark Aims to Reduce Number of False Positives Found by AI Vulnerability Scanners

DevOps.com, Wednesday, September 30th, 2026

AWS's Deception Benchmark shows AI models catch most real vulnerabilities but flag 41-99% of safe code.

DevOps.com reports that AWS built the Deception Benchmark, containing 14,822 adversarially hardened samples in 16 languages across more than 70 CWE categories, after evaluating 12 models from five providers on whether AI can tell exploitable vulnerabilities from code that only looks risky.

Models identified up to 95% of real vulnerabilities but flagged 41% to 99% of safe code, and proof-of-exploit prompting cut false positives yet missed 7% to 44% of real flaws.

Environment-gated challenges, such as checking whether a Kubernetes network policy blocks an SSRF path, proved hardest, and no configuration kept both error types below 10%.

more →  ·  More from AWS →