Search Results

Documents authored by Carbonell, Lanaya A.


Document
Emerging Results, Vision & Reflection Track Paper
CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts

Authors: Lanaya A. Carbonell and Anıl Koyuncu

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Adding vulnerability context to a prompt is known to improve LLM repair quality, but the studies that establish this vary two things at once: the domain knowledge injected, such as a CWE label or a fix pattern, and the form that knowledge takes in the prompt. Prompt-engineering research treats format as a variable in its own right, yet in automated vulnerability repair it has never been isolated under a controlled ablation, leaving open whether the wording of guidance matters once its content is fixed. This paper presents CWEFT, a controlled ablation that holds injected repair knowledge constant across a free-text and a typed-schema rendering while decomposing the separate contributions of naming a vulnerability and supplying fix guidance. We construct four prompt conditions that differ by exactly one factor each, grounding all injected knowledge in APR4Vul’s empirically mined fix patterns. Using this design, we generate 504 patches across 42 Java vulnerabilities from Vul4J spanning 15 CWE classes on three frontier models - Claude Sonnet 4.5, GPT-5, and Gemini 2.5 Pro - and score each patch on a Proof-of-Vulnerability-aware hierarchy that separates code that compiles from code that closes the vulnerability and from code that does so without regression. Our emerging results invert a common assumption: naming the CWE class alone achieves the highest correct-fix rate at 19.8%, ahead of the 12.7% unenriched baseline (p = 0.049, exact McNemar), while both richer conditions that add explicit repair guidance fall back to 16.7%. Crucially, the free-text and typed-schema conditions are indistinguishable at 16.7% with perfectly symmetric paired discordance, and with only eight discordant pairs the design cannot rule out a smaller format effect. Across all conditions, repair rate tracks how concrete the vulnerability class is, falling to 3.1% for improper input validation. These initial findings suggest that for frontier models on common vulnerability classes, what a prompt identifies matters more than how elaborately that information is formatted - a direction that warrants higher-powered replication before it is treated as a design principle.

Cite as

Lanaya A. Carbonell and Anıl Koyuncu. CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 78:1-78:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{carbonell_et_al:LIPIcs.ESEM.2026.78,
  author =	{Carbonell, Lanaya A. and Koyuncu, An{\i}l},
  title =	{{CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{78:1--78:14},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.78},
  URN =		{urn:nbn:de:0030-drops-280469},
  doi =		{10.4230/LIPIcs.ESEM.2026.78},
  annote =	{Keywords: Automated vulnerability repair, Automated program repair, Large language models, Prompt engineering, Structured prompting, Software security, Vul4J}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail