<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:32:54Z</responseDate>
  <request identifier="28046" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:28046</identifier>
        <datestamp>2026-10-05T06:44:06Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts</dc:title>
          <dc:creator>Carbonell, Lanaya A.</dc:creator>
          <dc:creator>Koyuncu, Anıl</dc:creator>
          <dc:subject>Automated vulnerability repair</dc:subject>
          <dc:subject>Automated program repair</dc:subject>
          <dc:subject>Large language models</dc:subject>
          <dc:subject>Prompt engineering</dc:subject>
          <dc:subject>Structured prompting</dc:subject>
          <dc:subject>Software security</dc:subject>
          <dc:subject>Vul4J</dc:subject>
          <dc:description>Adding vulnerability context to a prompt is known to improve LLM repair quality, but the studies that establish this vary two things at once: the domain knowledge injected, such as a CWE label or a fix pattern, and the form that knowledge takes in the prompt. Prompt-engineering research treats format as a variable in its own right, yet in automated vulnerability repair it has never been isolated under a controlled ablation, leaving open whether the wording of guidance matters once its content is fixed. This paper presents CWEFT, a controlled ablation that holds injected repair knowledge constant across a free-text and a typed-schema rendering while decomposing the separate contributions of naming a vulnerability and supplying fix guidance. We construct four prompt conditions that differ by exactly one factor each, grounding all injected knowledge in APR4Vul’s empirically mined fix patterns. Using this design, we generate 504 patches across 42 Java vulnerabilities from Vul4J spanning 15 CWE classes on three frontier models - Claude Sonnet 4.5, GPT-5, and Gemini 2.5 Pro - and score each patch on a Proof-of-Vulnerability-aware hierarchy that separates code that compiles from code that closes the vulnerability and from code that does so without regression. Our emerging results invert a common assumption: naming the CWE class alone achieves the highest correct-fix rate at 19.8%, ahead of the 12.7% unenriched baseline (p = 0.049, exact McNemar), while both richer conditions that add explicit repair guidance fall back to 16.7%. Crucially, the free-text and typed-schema conditions are indistinguishable at 16.7% with perfectly symmetric paired discordance, and with only eight discordant pairs the design cannot rule out a smaller format effect. Across all conditions, repair rate tracks how concrete the vulnerability class is, falling to 3.1% for improper input validation. These initial findings suggest that for frontier models on common vulnerability classes, what a prompt identifies matters more than how elaborately that information is formatted - a direction that warrants higher-powered replication before it is treated as a design principle.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Lanaya A. Carbonell and Anıl Koyuncu</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.78</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-280469</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.78</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
