What is it about?

Software teams use automated security scanners to find weaknesses in their code. These tools are useful, but they raise a flood of false alarms: warnings about code that is actually safe. Developers must check each warning by hand, which wastes time and leads many to ignore the tools altogether. We tested whether AI agents can do this checking for them. Unlike a chatbot that answers in one go, an AI agent can open files, follow the code across a project and test its own guesses before it decides. We compared three popular agent systems (Aider, OpenHands and SWE-agent), each paired with three AI models, on a standard Java test suite and on warnings from real Java projects. We then tried the best combination on recent C/C++ code that the AI models could not have seen during training.

Featured Image

Why is it important?

Security scanners are built into many development workflows, and AI agents are a new, fast-moving technology. This is the first study to compare agent systems head to head for checking scanner warnings. The best setup cut the share of safe code wrongly flagged from about 98% to about 6% on the standard test suite. On the unseen C/C++ code it caught 95.5% of the false alarms, against 36.4% for a plain one-shot prompt. So agents can take much of this tedious checking off developers' hands. The results also show where to be careful. Strong AI models gained a lot from working as agents, while for a weaker model the agent brought no consistent benefit. The agents were reliable for classic attacks such as code injection, but they wrongly dismissed many real problems involving weak encryption or security settings. Even the best setup dismissed more than one in five real vulnerabilities in the test suite. Costs also varied widely between agent systems. Our advice: use AI agents to help people decide, not to delete warnings automatically.

Perspectives

What surprised me most was how much the choice of AI model mattered compared with the choice of agent system. Agents amplified what a strong model could already do, but they did not make up for a weaker one. The finding I care about most is the trade-off. The same agent that clears nearly all false alarms for injection bugs also dismisses most real weak-cryptography issues, often treating them as best-practice advice rather than real risks. For me that settles how these tools should be used: as a second pair of eyes for developers, with the riskiest categories always going to a human. We have released all our scripts and data, including detailed logs of the agent runs, so others can study how these agents reason and not just what they conclude.

Nemo Xiong
Monash University

Read the Original

This page is a summary of: Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering, Proceedings of the ACM on Software Engineering, October 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3832100.
You can read the full text:

Read

Contributors

The following have contributed to this page