<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:44:54Z</responseDate>
  <request identifier="28008" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:28008</identifier>
        <datestamp>2026-10-05T06:44:04Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency</dc:title>
          <dc:creator>Liu, Junchi</dc:creator>
          <dc:creator>Bigdeli, Ali</dc:creator>
          <dc:creator>Daneshi, Roya</dc:creator>
          <dc:creator>Ambala, Atu</dc:creator>
          <dc:creator>Ghosh, Sudipto</dc:creator>
          <dc:creator>Santos, Fabio</dc:creator>
          <dc:subject>Automated Program Repair</dc:subject>
          <dc:subject>Large Language Models</dc:subject>
          <dc:subject>Fault Localization</dc:subject>
          <dc:subject>Bug Complexity</dc:subject>
          <dc:subject>Cost-efficiency</dc:subject>
          <dc:description>Background. Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on overall repair effectiveness. The effects of bug complexity, fault localization, reasoning settings, and repair cost-effectiveness remain insufficiently explored.&#13;
&#13;
Aims. This study presents a comprehensive empirical analysis of LLM-based APR, focusing on how repair performance is shaped by bug complexity, fault localization, reasoning settings, and costs.&#13;
&#13;
Method. We construct a curated dataset by collecting algorithmic bugs from AtCoder, a competitive programming platform. We evaluate two APR techniques (ChatRepair and CodeCorrector) using three LLMs (DeepSeek, GPT, and Llama), with multiple model variants and reasoning settings, and examine their performance across diverse levels of bug complexity and localization strategies through a multi-dimensional empirical framework and statistical analysis.&#13;
&#13;
Results. Although structurally complex bugs and imprecise fault localization make repair more challenging, LLM-based APR techniques still achieve competitive repair effectiveness. Imprecise fault localization can substantially enlarge the performance gap between APR techniques. Furthermore, higher-cost LLMs and stronger reasoning settings do not consistently yield better cost-efficiency, revealing a nontrivial trade-off between repair effectiveness and computational cost. We further observe that the impact of reasoning strategies varies considerably across different LLM families, affecting both repair effectiveness and cost-efficiency.&#13;
&#13;
Conclusions. Over 50% of moderately complex bugs can be repaired by low-cost LLM-based APR techniques. The repair effectiveness gap between APR techniques becomes larger as fault localization becomes less precise. GPT-5 repairs 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2, respectively; whereas the total repair cost of DeepSeek-V3.2 across the non-reasoning and reasoning stages is only approximately 4.6% and 7.7% of GPT-5 and DeepSeek-V4-pro, respectively.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Junchi Liu and Ali Bigdeli and Roya Daneshi and Atu Ambala and Sudipto Ghosh and Fabio Santos</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.40</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-280081</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.40</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
