<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:32:51Z</responseDate>
  <request identifier="28045" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:28045</identifier>
        <datestamp>2026-10-05T06:44:06Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR</dc:title>
          <dc:creator>Yengejeh, Sahand Moslemi</dc:creator>
          <dc:creator>Koyuncu, Anil</dc:creator>
          <dc:subject>automated program repair</dc:subject>
          <dc:subject>patch correctness assessment</dc:subject>
          <dc:subject>patch overfitting</dc:subject>
          <dc:description>Automated program repair (APR) tools generate candidate patches and accept those that pass a given test suite. A test-passing patch, called a plausible patch, is not necessarily correct: it may overfit the test suite without fixing the underlying bug. Automated patch correctness assessment (APCA) tools were developed to detect such overfitting patches, but they have been evaluated only as post-hoc classifiers on static datasets, never inside the repair loop they are intended to support. We investigate whether placing an APCA tool inside the generate-and-validate loop of an APR system, as a second validation gate, affects the APCA’s classification effectiveness, the APR tool’s repair capability, and the search cost it incurs. In this design, a candidate that passes all tests but is classified as overfitting by the APCA tool is discarded, and the repair tool continues generating candidates. We study it with three APR tools (TBar, ARJA, ChatRepair) crossed with three APCA tools (ODS, Quatrain, LLM4PatchCorrect), evaluated on 549 single-method bugs of Defects4J v3 with manual correctness assessment of 1,328 patches drawn from the in-loop candidate streams, and paired non-parametric tests (McNemar, Wilcoxon) per (APR, APCA) cell. Average in-loop F₁ on the correct class sits between 0.43 and 0.66 across the nine cells, with classification harder in-loop than post-hoc on every tool. The in-loop gate raises the share of correct outputs in 14 of 18 cells, but cuts the absolute correct-fix count whenever the APR baseline’s first plausible patch is already correct at a high rate. The in-loop introduces a search-cost overhead that is statistically significant in 14 of 18 cells (Wilcoxon p &lt; .05), with a geometric-mean per-bug multiplier of up to 17.46×. Placing an APCA inside the repair loop raises the output correctness rate in most (APR, APCA) cells, but at a search-cost overhead and, where the APR baseline already produces correct first plausibles at a high rate, a reduction in absolute correct fixes. The findings indicate that APCA tools should be trained and evaluated on the in-loop patch distribution they are deployed against, rather than on the curated, terminal patches of static post-hoc datasets.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Sahand Moslemi Yengejeh and Anil Koyuncu</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.77</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-280458</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.77</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
