Search Results

Documents authored by Santos, Fabio


Document
Technical Track Paper
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency

Authors: Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, and Fabio Santos

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on overall repair effectiveness. The effects of bug complexity, fault localization, reasoning settings, and repair cost-effectiveness remain insufficiently explored. Aims. This study presents a comprehensive empirical analysis of LLM-based APR, focusing on how repair performance is shaped by bug complexity, fault localization, reasoning settings, and costs. Method. We construct a curated dataset by collecting algorithmic bugs from AtCoder, a competitive programming platform. We evaluate two APR techniques (ChatRepair and CodeCorrector) using three LLMs (DeepSeek, GPT, and Llama), with multiple model variants and reasoning settings, and examine their performance across diverse levels of bug complexity and localization strategies through a multi-dimensional empirical framework and statistical analysis. Results. Although structurally complex bugs and imprecise fault localization make repair more challenging, LLM-based APR techniques still achieve competitive repair effectiveness. Imprecise fault localization can substantially enlarge the performance gap between APR techniques. Furthermore, higher-cost LLMs and stronger reasoning settings do not consistently yield better cost-efficiency, revealing a nontrivial trade-off between repair effectiveness and computational cost. We further observe that the impact of reasoning strategies varies considerably across different LLM families, affecting both repair effectiveness and cost-efficiency. Conclusions. Over 50% of moderately complex bugs can be repaired by low-cost LLM-based APR techniques. The repair effectiveness gap between APR techniques becomes larger as fault localization becomes less precise. GPT-5 repairs 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2, respectively; whereas the total repair cost of DeepSeek-V3.2 across the non-reasoning and reasoning stages is only approximately 4.6% and 7.7% of GPT-5 and DeepSeek-V4-pro, respectively.

Cite as

Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, and Fabio Santos. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 40:1-40:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{liu_et_al:LIPIcs.ESEM.2026.40,
  author =	{Liu, Junchi and Bigdeli, Ali and Daneshi, Roya and Ambala, Atu and Ghosh, Sudipto and Santos, Fabio},
  title =	{{Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{40:1--40:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.40},
  URN =		{urn:nbn:de:0030-drops-280081},
  doi =		{10.4230/LIPIcs.ESEM.2026.40},
  annote =	{Keywords: Automated Program Repair, Large Language Models, Fault Localization, Bug Complexity, Cost-efficiency}
}
Document
Technical Track Paper
Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-As-Code Repair

Authors: Benjamin Agyekum and Fabio Santos

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Iterative feedback loops have become the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed error signals back to the model for successive repair attempts. Prior work reports cumulative-best metrics, which are monotonically non-decreasing by construction, so the raw per-iteration security trajectory has never been examined in the IaC domain. Aims. We study security regression (a previously-passing CIS Benchmark check that fails after a repair iteration) to determine whether, and how often, iterative LLM repair degrades security while fixing other issues. Method. We analyze 5,968 scenario timelines from the IaC-Eval benchmark, each one scenario run through one configuration for up to 5 repair iterations. The 15 configurations comprise six model-specific RAG and nine model-aggregated non-RAG configurations, three temperatures each, and together they yield 4,440 iteration transitions with Checkov data on both sides. We track 30 individual CIS check IDs and classify regression root causes from code diffs, under two detection modes: standard (inclusive) and strict (exclusive check failures only). Results. Under standard (inclusive) detection, 13.8% of scenarios (24.8% of transitions) exhibit at least one regression. Under strict detection, which counts only unambiguous, exclusive check failures, the rate falls to 3.3% of scenarios (5.2% of transitions). This gap indicates that most apparent regressions are multi-resource measurement artifacts rather than genuine exclusive failures. Resource restructuring (79.0%) is the dominant root cause. Regression transitions show 2.6× more code churn (Cohen’s d = 0.90) and 4.9× higher strict-mode check volatility (d = 1.49). Of standard-mode regressions, 36.6% self-correct within an average of 1.2 iterations, and iteration 3 is the optimal stopping point. Conclusions. Iterative IaC repair does introduce security regressions, but most apparent regressions are multi-resource measurement artifacts. The conservative, defensible rate is approximately 3.3% of scenarios. Our findings motivate security-aware feedback-loop design and provide actionable iteration-budget guidance.

Cite as

Benjamin Agyekum and Fabio Santos. Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-As-Code Repair. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 50:1-50:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{agyekum_et_al:LIPIcs.ESEM.2026.50,
  author =	{Agyekum, Benjamin and Santos, Fabio},
  title =	{{Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-As-Code Repair}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{50:1--50:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.50},
  URN =		{urn:nbn:de:0030-drops-280186},
  doi =		{10.4230/LIPIcs.ESEM.2026.50},
  annote =	{Keywords: Infrastructure as Code, Security Regression, LLM Code Repair, Terraform, CIS Compliance, Iterative Feedback}
}
Document
Emerging Results, Vision & Reflection Track Paper
Towards Competence-Based Management for Open Source Software Projects

Authors: Sabahat Younas, Márcia Moraes, and Fabio Santos

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Contributors to Open Source Software (OSS) projects are vital to maintaining the health of both communities and projects. However, the number of projects experiencing core contributors' disengagement has increased to the point that it risks the projects' survival. Finding new contributors to replace the workforce is challenging and time-intensive due to several factors, including a lack of precise knowledge about potential candidates' competences, which may need to be confirmed through interviews and exams. Previous studies provided indications of the contributor’s competences, although they lack depth in understanding competence levels, which can result in poor knowledge about the contributor’s capabilities. To address this gap, we assess contributors' competence by collecting code metrics related to source code from contributions. We also propose a competence model able to predict the competence level required to solve tasks. By properly assessing contributors' competences, we can identify and train candidates to replace core contributors. Our replication package, including code, data, and documentation, is available at https://doi.org/10.5281/zenodo.21605122

Cite as

Sabahat Younas, Márcia Moraes, and Fabio Santos. Towards Competence-Based Management for Open Source Software Projects. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 65:1-65:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{younas_et_al:LIPIcs.ESEM.2026.65,
  author =	{Younas, Sabahat and Moraes, M\'{a}rcia and Santos, Fabio},
  title =	{{Towards Competence-Based Management for Open Source Software Projects}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{65:1--65:15},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.65},
  URN =		{urn:nbn:de:0030-drops-280339},
  doi =		{10.4230/LIPIcs.ESEM.2026.65},
  annote =	{Keywords: Contributor competence, MSR, project sustainability, AST, OSS, disengagement mitigation}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail