Search Results

Documents authored by Pan, Zulie


Document
Technical Track Paper
Do LLM Vulnerability-Detection Agents Reason Better? An Empirical Decomposition of Robustness, Resistance, and Grounding

Authors: Ziyu Chen, Qiangpu Chen, Yuwei Li, Taiyan Wang, Shiwen Ou, Qingsong Xie, Lu Zhang, and Zulie Pan

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Agents have recently improved LLM-based vulnerability detection through retrieval, iterative analysis, and verifier modules. However, existing evaluations mainly report improvements in detection performance without determining whether these improvements reflect more reliable vulnerability reasoning or access to retrieved context. Aims. This study investigates how different agent components affect the reasoning-oriented evaluation dimensions of agent-based vulnerability detection. Method. We present a controlled measurement framework with six conditions and evaluate it on six base models. The conditions progress from plain single-pass prompting through matched-context replay and ReAct-style tool use to self-review, trace-isolated verification, and multi-role decomposition, allowing us to separate context effects, iterative orchestration, and added agent modules under a shared protocol. Beyond vulnerability detection performance metrics, our framework evaluates three reasoning-oriented dimensions: robustness under semantics-preserving perturbations, resistance to misleading non-causal context, and grounding in concrete code evidence. Results. The full agent often achieves higher accuracy or F1 than plain single-pass prompting, but these improvements do not consistently extend to the three measured behavioral dimensions of robustness, resistance, and grounding. The models are not consistently more stable under harmless code changes, are often misled by non-causal prompt cues, and rarely cite the exact code lines that explain the vulnerability. In the strongest improvement case, much of the accuracy change already appears when the model receives the same retrieved tool outputs without the full agent process; after that context is held fixed, the remaining agent effect is mainly associated with predicting vulnerable more often. Conclusions. Higher detection performance should not be taken as direct evidence of better vulnerability reasoning. In our experiments, better performance metrics can reflect retrieved context or changes in how often the model predicts vulnerable, rather than a clear improvement in stable, evidence-based vulnerability reasoning. Future evaluations should report these behavior checks alongside performance metrics.

Cite as

Ziyu Chen, Qiangpu Chen, Yuwei Li, Taiyan Wang, Shiwen Ou, Qingsong Xie, Lu Zhang, and Zulie Pan. Do LLM Vulnerability-Detection Agents Reason Better? An Empirical Decomposition of Robustness, Resistance, and Grounding. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 28:1-28:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{chen_et_al:LIPIcs.ESEM.2026.28,
  author =	{Chen, Ziyu and Chen, Qiangpu and Li, Yuwei and Wang, Taiyan and Ou, Shiwen and Xie, Qingsong and Zhang, Lu and Pan, Zulie},
  title =	{{Do LLM Vulnerability-Detection Agents Reason Better? An Empirical Decomposition of Robustness, Resistance, and Grounding}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{28:1--28:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.28},
  URN =		{urn:nbn:de:0030-drops-279961},
  doi =		{10.4230/LIPIcs.ESEM.2026.28},
  annote =	{Keywords: Large language models, vulnerability detection, software security, agent-based systems, empirical evaluation}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail