Search Results

Documents authored by Koyuncu, Anil


Document
Emerging Results, Vision & Reflection Track Paper
APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR

Authors: Sahand Moslemi Yengejeh and Anil Koyuncu

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Automated program repair (APR) tools generate candidate patches and accept those that pass a given test suite. A test-passing patch, called a plausible patch, is not necessarily correct: it may overfit the test suite without fixing the underlying bug. Automated patch correctness assessment (APCA) tools were developed to detect such overfitting patches, but they have been evaluated only as post-hoc classifiers on static datasets, never inside the repair loop they are intended to support. We investigate whether placing an APCA tool inside the generate-and-validate loop of an APR system, as a second validation gate, affects the APCA’s classification effectiveness, the APR tool’s repair capability, and the search cost it incurs. In this design, a candidate that passes all tests but is classified as overfitting by the APCA tool is discarded, and the repair tool continues generating candidates. We study it with three APR tools (TBar, ARJA, ChatRepair) crossed with three APCA tools (ODS, Quatrain, LLM4PatchCorrect), evaluated on 549 single-method bugs of Defects4J v3 with manual correctness assessment of 1,328 patches drawn from the in-loop candidate streams, and paired non-parametric tests (McNemar, Wilcoxon) per (APR, APCA) cell. Average in-loop F₁ on the correct class sits between 0.43 and 0.66 across the nine cells, with classification harder in-loop than post-hoc on every tool. The in-loop gate raises the share of correct outputs in 14 of 18 cells, but cuts the absolute correct-fix count whenever the APR baseline’s first plausible patch is already correct at a high rate. The in-loop introduces a search-cost overhead that is statistically significant in 14 of 18 cells (Wilcoxon p < .05), with a geometric-mean per-bug multiplier of up to 17.46×. Placing an APCA inside the repair loop raises the output correctness rate in most (APR, APCA) cells, but at a search-cost overhead and, where the APR baseline already produces correct first plausibles at a high rate, a reduction in absolute correct fixes. The findings indicate that APCA tools should be trained and evaluated on the in-loop patch distribution they are deployed against, rather than on the curated, terminal patches of static post-hoc datasets.

Cite as

Sahand Moslemi Yengejeh and Anil Koyuncu. APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 77:1-77:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{yengejeh_et_al:LIPIcs.ESEM.2026.77,
  author =	{Yengejeh, Sahand Moslemi and Koyuncu, Anil},
  title =	{{APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{77:1--77:15},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.77},
  URN =		{urn:nbn:de:0030-drops-280458},
  doi =		{10.4230/LIPIcs.ESEM.2026.77},
  annote =	{Keywords: automated program repair, patch correctness assessment, patch overfitting}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail