LLMScanBench Q3 2026: Choosing the Right Approach for AI Vulnerability Discovery
Harness, Monday, September 28th, 2026
Harness benchmark: LLM vulnerability scanners show high precision but weak recall and run slower and costlier than AI SAST.
Harness's LLMScanBench Q3 2026 compares LLM-based vulnerability scanners with its AI SAST on precision, recall, speed, cost and consistency.
LLM scanners showed high precision but fell short on recall, with GPT 5.5 outperforming Opus 4.8 in most cases.
Because LLM scanners are probabilistic, repeated scans of unchanged code produced different findings, token use and run times, and they took 1.6-37x longer and cost 6-39x more than Harness AI SAST.