No documents found matching your filter selection.
Document
Complete Volume
Authors:
Robert Feldt, Maria Paasivaara, Daniel Mendez, Stefan Wagner, and Marvin Muñoz Barón
Abstract
LIPIcs, Volume 394, ESEM 2026, Complete Volume
Cite as
20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 1-1786, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@Proceedings{feldt_et_al:LIPIcs.ESEM.2026,
title = {{LIPIcs, Volume 394, ESEM 2026, Complete Volume}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {1--1786},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026},
URN = {urn:nbn:de:0030-drops-280981},
doi = {10.4230/LIPIcs.ESEM.2026},
annote = {Keywords: LIPIcs, Volume 394, ESEM 2026, Complete Volume}
}
Document
Front Matter
Authors:
Robert Feldt, Maria Paasivaara, Daniel Mendez, Stefan Wagner, and Marvin Muñoz Barón
Abstract
Front Matter, Table of Contents, Preface, Conference Organization
Cite as
20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 0:i-0:xxvi, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{feldt_et_al:LIPIcs.ESEM.2026.0,
author = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
title = {{Front Matter, Table of Contents, Preface, Conference Organization}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {0:i--0:xxvi},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.0},
URN = {urn:nbn:de:0030-drops-280960},
doi = {10.4230/LIPIcs.ESEM.2026.0},
annote = {Keywords: Front Matter, Table of Contents, Preface, Conference Organization}
}
Document
Technical Track Paper
Authors:
Quanzhi Fu, Wang Lingxiang, Wenjia Song, Gelei Deng, Yi Liu, Dan Williams, and Ying Zhang
Abstract
Background. The integration of open-source libraries in Java development introduces severe security risks through vulnerable APIs. Existing program analysis and deep learning tools face challenges in capturing inter-procedural vulnerability semantics at scale. While LLMs show promise for semantic reasoning, they cannot handle large codebases due to context limits, and they lack the vulnerability-specific understanding needed to determine exploitability.
Aim. This work aims to overcome these limitations and enable LLM-based detection of vulnerable API usage in large-scale Java applications.
Method. We present CognixShield, an LLM-powered framework for detecting vulnerable API usage through three core components. First, semantic-preserving AST-based fragmentation partitions large codebases while maintaining syntactic completeness within LLM context windows. Second, vulnerability-aware multi-agent RAG traces relevant program context across these fragments, iteratively assembling security-critical context spanning functions and files. Third, PoV-guided semantic reasoning uses Proof-of-Vulnerability tests that encode precise triggering conditions and exploitation mechanics to determine vulnerability.
Results. CognixShield achieves 84% precision, 95% recall, 84% accuracy, and an 89% F1-score on 57 real-world Java applications, outperforming state-of-the-art tools.
Conclusions. Our results show that vulnerability detection requires specialized architectural innovations beyond generic LLM applications.
Cite as
Quanzhi Fu, Wang Lingxiang, Wenjia Song, Gelei Deng, Yi Liu, Dan Williams, and Ying Zhang. CognixShield: PoV-Guided Vulnerable API Usage Detection in Large Codebases via LLMs. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 1:1-1:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{fu_et_al:LIPIcs.ESEM.2026.1,
author = {Fu, Quanzhi and Lingxiang, Wang and Song, Wenjia and Deng, Gelei and Liu, Yi and Williams, Dan and Zhang, Ying},
title = {{CognixShield: PoV-Guided Vulnerable API Usage Detection in Large Codebases via LLMs}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {1:1--1:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.1},
URN = {urn:nbn:de:0030-drops-279698},
doi = {10.4230/LIPIcs.ESEM.2026.1},
annote = {Keywords: Vulnerable API usage detection, program analysis, LLMs, agentic RAG}
}
Document
Technical Track Paper
Authors:
Alexandros Tsakpinis, Emil Schwenger, and Alexander Pretschner
Abstract
Background. Open source software ecosystems exhibit dense dependency networks in which maintenance degradation of structurally central packages can propagate widely. Despite increasing attention to open source sustainability, existing support mechanisms lack an explicit, dependency-aware notion of ecosystem-level impact to guide support decisions.
Aims. In this paper, we introduce a dependency-aware model of ecosystem impact that captures how changes in maintenance activities propagate through the Python Package Index (PyPI) ecosystem and affect its overall state. Based on this model, we prioritize packages for ecosystem support using our dependency-propagated notion of ecosystem impact.
Method. Applying this framework to a snapshot of 718,750 PyPI packages and over 2 million dependencies, we compare our impact-driven support strategy with existing support mechanisms (Tidelift, Ecosyste.ms, and GitHub Sponsors) and with PageRank as a baseline measure of structural importance.
Results. Our results show that a large share of the modeled ecosystem impact (approximately 80%) can be attributed to just 0.1% of all PyPI packages when prioritized based on dependency-propagated impact. In contrast, externally defined support sets vary substantially in their alignment with ecosystem impact. We further analyze maintainer reach and metadata accessibility, revealing that ecosystem impact, social footprint, and operational feasibility represent distinct but complementary dimensions of ecosystem support.
Conclusions. Dependency-aware ecosystem impact modeling provides a transparent and systematic basis for prioritizing support in large-scale software ecosystems. Our findings suggest that effective support strategies, driven by ecosystem stewards, funding bodies, and organizations operating support programs, should complement existing allocation logic with impact-informed decision making.
Cite as
Alexandros Tsakpinis, Emil Schwenger, and Alexander Pretschner. Modeling Dependency-Propagated Ecosystem Impact of Changes in Maintenance Activities: Evaluating Support Strategies in the PyPI Network. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 2:1-2:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{tsakpinis_et_al:LIPIcs.ESEM.2026.2,
author = {Tsakpinis, Alexandros and Schwenger, Emil and Pretschner, Alexander},
title = {{Modeling Dependency-Propagated Ecosystem Impact of Changes in Maintenance Activities: Evaluating Support Strategies in the PyPI Network}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {2:1--2:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.2},
URN = {urn:nbn:de:0030-drops-279708},
doi = {10.4230/LIPIcs.ESEM.2026.2},
annote = {Keywords: OSS Ecosystems, Dependency Networks, Ecosystem Impact, OSS Support}
}
Document
Technical Track Paper
Authors:
Ebtesam Al-Haque and Brittany Johnson
Abstract
Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and characterization of task difficulty. More specifically, we currently have little understanding of what makes one task harder than another, and to what extent task difficulty is predictable from static task properties.
Aims. We propose a measurement framework to investigate and systematically quantify what structural properties of software tasks correspond to agent success rates for issue resolution tasks.
Method. We conducted a large scale empirical study on CoderForge-Preview, the largest open dataset of coding agent trajectories to date, by extracting features across task patch, repository and prompt. We evaluated the predictive power of each feature against task outcomes using ensemble methods, SHAP attribution, and effect size analysis.
Results. We found that task difficulty is substantially predictable from static features (AUC = 0.863) and is largely driven by patch fragmentation and repository scale. Prompt linguistic features become visible among top contributors for tasks in the mid-band, revealing a layered structure of difficulty.
Conclusion. The difficulty of an issue resolution task is encoded in its structure. This enables static, pre-hoc difficulty estimation and lays the groundwork for difficulty-controlled benchmark construction for evaluation of agents.
Cite as
Ebtesam Al-Haque and Brittany Johnson. What Makes Software Issue Resolution Tasks Difficult for Agents?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 3:1-3:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{alhaque_et_al:LIPIcs.ESEM.2026.3,
author = {Al-Haque, Ebtesam and Johnson, Brittany},
title = {{What Makes Software Issue Resolution Tasks Difficult for Agents?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {3:1--3:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.3},
URN = {urn:nbn:de:0030-drops-279712},
doi = {10.4230/LIPIcs.ESEM.2026.3},
annote = {Keywords: software, oss, agents, empirical studies}
}
Document
Technical Track Paper
Authors:
Michel Albonico and Nuno Macedo
Abstract
Background. The Robot Operating System (ROS) is a framework widely adopted for developing robotics applications. Its topic-based communication model allows distributed components to exchange data via message-passing, where suboptimal messaging configuration can lead to performance degradation and resource inefficiency.
Aims. This paper aims to investigate the energy consumption for ROS 2 topic-based communication, focusing on messaging configurations commonly used in real-world robotics projects.
Method. We systematically mine 112 active ROS 2 GitHub repositories to identify prevalent message types, frequencies, sizes, and number of subscribers. These insights guide the design of a controlled experiment in which we measure the power consumption of ROS 2 nodes across multiple configurations.
Results. Analysis reveals that message type and interval significantly affect subscriber-side energy usage, with complex messages also affecting the publisher side. This evidence helps guide the design of energy-aware robotics software based on publish/subscribe communication.
Conclusions. ROS 2 topic-based communication design decisions, especially message type and publishing interval, have a substantial impact on energy consumption.
Cite as
Michel Albonico and Nuno Macedo. Whispers and Watts: An Empirical Study of the Energy Consumption of ROS 2 Topic-Based Communication. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 4:1-4:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{albonico_et_al:LIPIcs.ESEM.2026.4,
author = {Albonico, Michel and Macedo, Nuno},
title = {{Whispers and Watts: An Empirical Study of the Energy Consumption of ROS 2 Topic-Based Communication}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {4:1--4:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.4},
URN = {urn:nbn:de:0030-drops-279723},
doi = {10.4230/LIPIcs.ESEM.2026.4},
annote = {Keywords: ROS 2, Energy Consumption, Topic Communication, GitHub Mining}
}
Document
Technical Track Paper
Authors:
Xingcheng Chen, Mehmet Besenk, and Andrea Stocco
Abstract
Background. Task-specialized language models are increasingly integrated into software engineering workflows to support vertical-domain activities such as issue triaging, document classification, and automated analysis. Despite their adoption, there is limited empirical evidence on how to test their robustness and detect brittle behaviors under semantics-preserving input transformations.
Aims. This paper investigates whether explainability-guided metamorphic testing can improve the effectiveness and validity of robustness testing for specialized language models compared to heuristic mutation strategies.
Method. We conduct a large-scale empirical study of explanation-guided metamorphic testing across three datasets, four model architectures, and 20 testing configurations derived from combinations of attribution methods and mutation strategies. The evaluated configurations combine attribution-based token prioritization, LLM-driven mutation, and automated semantic verification to generate linguistically valid test variants. We assess failure discovery capability, semantic validity, and testing efficiency against heuristic baselines.
Results. Explanation-guided metamorphic testing generates 2.30× more verified failure-inducing test cases than heuristic mutation strategies. Semantic verification substantially improves mutation validity and achieves high label-preservation precision among gate-accepted variants according to human annotation. The study further reveals systematic shortcut behaviors across models, including over-reliance on named entities and formatting cues.
Conclusions. The results provide evidence that explanation-guided metamorphic testing is an effective and practical approach for empirically evaluating the robustness of task-specialized language models used in vertical AI applications.
Cite as
Xingcheng Chen, Mehmet Besenk, and Andrea Stocco. Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 5:1-5:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{chen_et_al:LIPIcs.ESEM.2026.5,
author = {Chen, Xingcheng and Besenk, Mehmet and Stocco, Andrea},
title = {{Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {5:1--5:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.5},
URN = {urn:nbn:de:0030-drops-279739},
doi = {10.4230/LIPIcs.ESEM.2026.5},
annote = {Keywords: Metamorphic testing, explainable AI, specialized language models, vertical AI, attribution-guided testing, automated test generation, empirical software engineering}
}
Document
Technical Track Paper
Authors:
Ronnie de Souza Santos, Italo Santos, Maria Teresa Baldassarre, Cleyton Magalhães, and Mairieli Wessel
Abstract
Background. Large Language Models (LLMs) introduce new concerns regarding fraudulent or AI-assisted participation in software engineering surveys.
Aims. This study investigates how suspicious or potentially AI-assisted responses may affect the validity of software engineering survey findings.
Method. We conducted a secondary analysis of four software engineering survey datasets using manual identification of suspicious responses, automated AI-generated text detection, descriptive statistical analysis, and thematic analysis. We compared findings obtained from the original and manually cleaned datasets.
Results. Quantitative findings generally remained stable after filtering suspicious responses, although some demographic and analytical variables showed moderate variation, affecting the interpretation of specific participant groups and contextual characteristics. In contrast, qualitative findings were more strongly influenced by changes in contextual framing, code prominence, and the nature of the evidence supporting interpretation, shaping how participants' experiences and study contexts were interpreted and characterized.
Conclusions. AI-assisted participation may influence software engineering survey findings differently depending on the type of analysis being conducted. The findings reinforce the importance of combining multiple validation procedures, particularly in studies relying on open-ended responses.
Cite as
Ronnie de Souza Santos, Italo Santos, Maria Teresa Baldassarre, Cleyton Magalhães, and Mairieli Wessel. The Influence of Fraudulent AI-Generated Responses on Software Engineering Surveys. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 6:1-6:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{desouzasantos_et_al:LIPIcs.ESEM.2026.6,
author = {de Souza Santos, Ronnie and Santos, Italo and Baldassarre, Maria Teresa and Magalh\~{a}es, Cleyton and Wessel, Mairieli},
title = {{The Influence of Fraudulent AI-Generated Responses on Software Engineering Surveys}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {6:1--6:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.6},
URN = {urn:nbn:de:0030-drops-279748},
doi = {10.4230/LIPIcs.ESEM.2026.6},
annote = {Keywords: LLMs, survey, threats to validity}
}
Document
Technical Track Paper
Authors:
Nitish Patkar, Xeno Isenegger, Gideon Monterosa, Sebastiano Panichella, and Norbert Seyff
Abstract
Background. Meetings are an integral part of software practitioners' workday, but many practitioners see them as disruptive to focused work and productivity. Prior research has mainly focused on overall meeting load and broad perceptions; however, teams cannot effectively improve meeting practices without knowing which specific meeting patterns cause disruption.
Aims. Our core aim is to move from general perceptions to an empirical investigation of meeting characteristics that are strongly linked to productivity, flow, and well-being.
Method. We conducted a two-phase mixed-method study. In Phase 1, we surveyed 55 software practitioners to establish a baseline of meeting practices and perceived impact on productivity, flow, and well-being. In Phase 2, a two-week field study with experience sampling collected real-time post-meeting ratings (usefulness, energy, focus disruption) and daily end-of-day reports, complemented by objective meeting characteristics extracted automatically from calendar metadata (135 post-meeting and 40 end-of-day records).
Results. In the baseline survey, 61.8% reported at least moderate flow disruption and 76.4% lower productivity on high-meeting days, with unclear purpose and excessive duration as the top frustrations. The field study revealed a more nuanced pattern: overall usefulness was moderately positive, but varied by type - ad-hoc meetings were among the most useful (Mdn = 4), planning meetings were most disruptive (Mdn = 4), and one-on-ones were least disruptive (Mdn = 1). Type contrasts are exploratory and based on small sample. At the daily level, higher meeting counts coincided with lower productivity and deep-work ratings; energy was lowest on high-meeting days. This daily pattern reflected between-participant differences rather than within-person day-to-day effects. The directional alignment between survey complaints and real-time observations strengthens confidence in these emerging patterns.
Conclusions. This study shifts the conversation from broad dissatisfaction with meetings toward identifiable and actionable meeting patterns that organizations can actively redesign. The findings suggest that not all meetings are equally harmful-or beneficial-and that factors such as meeting purpose clarity, duration, and scheduling context may play a critical role in shaping employees’ productivity, ability to sustain flow, and daily energy levels. We treat these patterns as design hypotheses pending larger confirmation. They also provide an empirical starting point for longitudinal research on meeting design.
Cite as
Nitish Patkar, Xeno Isenegger, Gideon Monterosa, Sebastiano Panichella, and Norbert Seyff. Characterising Meeting Practices in Software Engineering Teams: A Mixed-Methods Study of Perceptions, Objective Patterns, and Impacts. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 7:1-7:19, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{patkar_et_al:LIPIcs.ESEM.2026.7,
author = {Patkar, Nitish and Isenegger, Xeno and Monterosa, Gideon and Panichella, Sebastiano and Seyff, Norbert},
title = {{Characterising Meeting Practices in Software Engineering Teams: A Mixed-Methods Study of Perceptions, Objective Patterns, and Impacts}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {7:1--7:19},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.7},
URN = {urn:nbn:de:0030-drops-279754},
doi = {10.4230/LIPIcs.ESEM.2026.7},
annote = {Keywords: Software engineering, Meeting practices, Productivity, Flow (deep work), Well-being, Experience sampling method (ESM)}
}
Document
Technical Track Paper
Authors:
Johan Linåker, Kevin Lumbard, and Georg Link
Abstract
Background. Open Source Software (OSS) remains underfunded, threatening the sustainability of critical digital infrastructure. In response, several public, private, and philanthropic funding initiatives are emerging. However, demand greatly exceeds supply, underscoring the need to increase funding and improve the efficiency and impact of existing funding sources.
Aims. This study investigates how OSS funders conceptualize, measure, and report funding impact, and the implications for sustaining and funding OSS projects. We focus on how funders select projects, define objectives, measure outcomes, and justify continued funding.
Method. We conducted a qualitative study based on semi-structured interviews with 25 participants representing funders, funding recipients, and research organizations in the OSS ecosystem. The findings were synthesized using thematic analysis and validated and enriched through a follow-up workshop with 16 participants from overlapping organizations.
Results. We find that impact is typically measured through short-term, milestone-based deliverables and qualitative narratives, and is dependent on personal funder-recipient relationships. Long-term sustainability, risk mitigation, and human-centered impacts are difficult to capture and rarely measured systematically. Measurement practices are shaped by (often implicit) funder strategies, funding structures, and reporting requirements, leading to ad-hoc and project-specific approaches.
Conclusions. Measuring and reporting funding impact remains a shared and non-trivial challenge for OSS funders. Addressing this requires explicit impact logics and shared meta‑frameworks (rather than universal metrics) to enable a more coordinated and complementary OSS funding ecosystem. This, in turn, can improve the effectiveness, accountability, and availability of funding and ultimately support the long-term sustainability of OSS as critical infrastructure.
Cite as
Johan Linåker, Kevin Lumbard, and Georg Link. Making Impact Visible in Funding of Open Source Software: A Study of the Funders' Perspective. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 8:1-8:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{linaker_et_al:LIPIcs.ESEM.2026.8,
author = {Lin\r{a}ker, Johan and Lumbard, Kevin and Link, Georg},
title = {{Making Impact Visible in Funding of Open Source Software: A Study of the Funders' Perspective}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {8:1--8:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.8},
URN = {urn:nbn:de:0030-drops-279769},
doi = {10.4230/LIPIcs.ESEM.2026.8},
annote = {Keywords: Open Source Software, Funding, Project Health, Community Health, Sustainability}
}
Document
Technical Track Paper
Authors:
Theocharis Tavantzis, Stefano Lambiase, and Daniel Russo
Abstract
Background. Generative AI (GenAI) makes software engineers faster, yet a growing number of them report feeling less like engineers. They describe feeling disconnected from their own code, uncertain about their skills, and uneasy about what their profession is becoming. Prior studies capture fragments of this experience in isolation, yet no shared concept or framework describes it as a whole.
Aim. Drawing on postphenomenology, Floridi’s re-ontologization of the infosphere, and Feenberg’s critical theory of technology, we introduce AI-induced displacement: the multi-dimensional way in which GenAI tools unsettle engineers' sense of craft, identity, and professional belonging.
Method. We interviewed 21 software engineers across different roles and organizations, and analyzed their accounts through thematic analysis.
Results. We identify six dimensions of displacement, concerning (i) how engineers justify and adopt AI-generated solutions, (ii) shifts in professional identity, (iii) tensions with normative and ethical standards, (iv) disruptions to individual focus and cognitive flow, (v) changing dynamics of recognition and collaboration within teams, and (vi) the erosion of intrinsic meaning and long-term professional sustainability.
Conclusions. This typology offers researchers a theoretically grounded lens for investigating the human side of GenAI adoption in software engineering, and helps practitioners design integration strategies that foreground augmentation rather than replacement.
Cite as
Theocharis Tavantzis, Stefano Lambiase, and Daniel Russo. Out of Place in My Own Work: How Generative AI Displaces Software Engineers. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 9:1-9:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{tavantzis_et_al:LIPIcs.ESEM.2026.9,
author = {Tavantzis, Theocharis and Lambiase, Stefano and Russo, Daniel},
title = {{Out of Place in My Own Work: How Generative AI Displaces Software Engineers}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {9:1--9:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.9},
URN = {urn:nbn:de:0030-drops-279777},
doi = {10.4230/LIPIcs.ESEM.2026.9},
annote = {Keywords: AI-induced displacement, Generative AI, Human-centered AI, Software Engineering}
}
Document
Technical Track Paper
Authors:
Hera Arif, Miikka Kuutila, and Paul Ralph
Abstract
Background. Code quality metrics are intended to measure latent properties of software source code. Although numerous code metrics have been proposed and used, their construct validity is rarely evaluated. Thus, the extent to which code metrics actually measure what they claim to measure is often unclear.
Aim. Drawing from modern measurement theory, we investigate the construct validity of common class-level, object-oriented code quality metrics.
Method. As code quality metrics are intended to reflect latent attributes, such as cohesion and coupling, we identified the factor structure of code quality metrics using Exploratory Factor Analysis (EFA). The metrics were extracted from the Apache Maven project by three software tools: Designite, JHawk, and Understand. The factor structure was later verified using Confirmatory Factor Analysis (CFA) on 22 randomly selected open source projects meeting a predetermined eligibility criteria.
Results. 24 code quality metrics that correspond to six constructs: Cohesion, In-Coupling, Out-Coupling, Size, Sub-Inheritance (related to subclasses), and Sup-Inheritance (related to superclasses) were revealed in the underlying factor structure. Ten metrics did not correspond to any known dimension of software quality and were removed in the exploratory analysis. Ten additional metrics exhibited low loadings in the confirmatory analysis, suggesting their removal from the final measurement model. Size, Cohesion, Inheritance, and Coupling were the constructs retained, with subcategories identified for Inheritance and Coupling.
Conclusions. Our results strongly support the construct validity of 24 code quality metrics. Coupling and Inheritance are revealed as multidimensional constructs, since they require measuring two different concepts, revealed as sub-categories in our analysis, and Complexity may be better explored in a multilevel model. Our results also corroborate the relationship between Cohesion, Size, and Out-Coupling which can be further explored in a structural model. Some metrics from the Chidamber & Kemerer metrics suite are found to perhaps be measuring different constructs than intended. Our results reveal the need for creating or integrating metrics that reflect existing constructs but measure fundamentally different properties to improve the content validity of the measurement model. Additionally, we provide useful recommendations for researchers, developers, and tool providers which stem from theoretical and empirical justification. Overall, our study demonstrates the value of applying modern measurement theory and latent variable modeling in validating software code quality metrics.
Cite as
Hera Arif, Miikka Kuutila, and Paul Ralph. Assessing the Construct Validity of Object-Oriented, Class-Level Code Quality Metrics. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 10:1-10:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{arif_et_al:LIPIcs.ESEM.2026.10,
author = {Arif, Hera and Kuutila, Miikka and Ralph, Paul},
title = {{Assessing the Construct Validity of Object-Oriented, Class-Level Code Quality Metrics}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {10:1--10:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.10},
URN = {urn:nbn:de:0030-drops-279781},
doi = {10.4230/LIPIcs.ESEM.2026.10},
annote = {Keywords: Code quality metrics, Factor analysis, Exploratory factor analysis, Confirmatory factor analysis, Software quality, Size, Inheritance, Coupling, Cohesion}
}
Document
Technical Track Paper
Authors:
Noah Leu, Julian Oertel, and Regina Hebig
Abstract
Background. Generating software tests and implementations with Large Language Models (LLMs) is becoming more common. However, using the same LLM for the generation of implementation and tests might make both suffer from the same biases, potentially decreasing the tests' ability to detect faults compared to tests generated by a different LLM.
Aims. In this paper, we want to investigate this notion by combining different LLMs for generating implementation and tests.
Method. We generate Elixir code and test cases with 5 LLMs and compare them to a baseline of manually written code and tests. We wrote and generated 459 distinct implementations and 2709 test suites, manually analyzing 9417 test case failures.
Results. We find no difference between tests generated by the same or a different LLM as the code. However, of the 247 implementations flagged as faulty, 99 were only flagged by manually written tests.
Conclusions. Our results indicate that LLM-generated tests alone are not yet sufficient.
Cite as
Noah Leu, Julian Oertel, and Regina Hebig. Don't Bother to Use a Second LLM and Write Tests Yourself! A Study on Elixir. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 11:1-11:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{leu_et_al:LIPIcs.ESEM.2026.11,
author = {Leu, Noah and Oertel, Julian and Hebig, Regina},
title = {{Don't Bother to Use a Second LLM and Write Tests Yourself! A Study on Elixir}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {11:1--11:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.11},
URN = {urn:nbn:de:0030-drops-279793},
doi = {10.4230/LIPIcs.ESEM.2026.11},
annote = {Keywords: LLMs, Code Generation, Test Generation, Rare Languages}
}
Document
Technical Track Paper
Authors:
Sergio Cobos, Robert Clarisó, and Javier Luis Cánovas Izquierdo
Abstract
Background. Large Language Models (LLMs) are increasingly used in software pipelines as decision components that validate, triage, or score system inputs and outputs. When one such component is replaced, for example, to reduce cost or latency or to adapt to provider-side model updates, it is necessary to make sure that the replacement does not significantly alter system behavior.
Aims. This paper presents a methodology for assessing the operational substitutability of LLM-based decision components under controlled conditions. We combine two complementary criteria: instance-level behavioral agreement, measured with Cohen’s κ, and the absence of statistically significant asymmetries in correctness outcomes, assessed with McNemar’s test.
Method. We instantiate the methodology on three representative binary decision tasks used in software pipelines: safety classification, truthfulness assessment, and stereotype detection. Using models from multiple providers, we study substitutability by asking (1) whether replacement within the same provider preserves decision behavior, (2) whether replacement across providers preserves decision behavior, and (3) how sensitive the admissible substitution space is to the agreement threshold and significance level used in the analysis.
Results. Substitutability varies substantially across tasks and model pairs. Safety admits the broadest substitution space, while truthfulness is the most constrained task within providers and stereotype detection across providers. Sensitivity analyses show that operational parameter choices change the size of the admissible substitution space, but not the overall qualitative pattern.
Conclusions. When assessing replacement candidates for LLM-based decision components, sufficient behavioral agreement and symmetry in correctness outcomes can complement aggregate benchmark performance.
Cite as
Sergio Cobos, Robert Clarisó, and Javier Luis Cánovas Izquierdo. Assessing the Substitution of LLM-Based Decision Components. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 12:1-12:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{cobos_et_al:LIPIcs.ESEM.2026.12,
author = {Cobos, Sergio and Claris\'{o}, Robert and C\'{a}novas Izquierdo, Javier Luis},
title = {{Assessing the Substitution of LLM-Based Decision Components}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {12:1--12:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.12},
URN = {urn:nbn:de:0030-drops-279803},
doi = {10.4230/LIPIcs.ESEM.2026.12},
annote = {Keywords: LLM-based decision components, model substitutability, LLM-as-a-judge, behavioral agreement, model replacement, software validation}
}
Document
Technical Track Paper
Authors:
Arumoy Shome, Luís Cruz, Diomidis Spinellis, and Arie van Deursen
Abstract
Background. Machine learning development in Jupyter notebooks is iterative and feedback-driven. Practitioners author statements that reveal information about program execution and use this information to decide what to do next. We call these feedback statements and identify two forms: exploratory statements that display values for visual inspection and validation statements that enforce conditions programmatically through assertions.
Aims. Many failures in ML systems do not surface as exceptions, and consequently escape the crash-based analyses that dominate prior empirical work on ML notebooks. This study examines what practitioners check to catch the failures that would otherwise pass silently, by characterizing feedback statements that encode the practitioner’s mental model of what the code should do and what could go wrong.
Method. We mine 297,851 publicly available Python Jupyter notebooks from GitHub and Kaggle, and extract 1,092,780 feedback statements. We sample 816 statements through proportional stratified sampling from semantic clusters obtained from CodeBERT embeddings, and apply grounded theory and open coding to manually label and analyze each statement.
Results. We contribute a taxonomy of feedback statements in ML Jupyter notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears. The taxonomy reveals that feedback in ML notebooks is overwhelmingly exploratory, and that the two platforms host qualitatively different modes of ML work. We further map our taxonomy to an existing crash taxonomy and find that our taxonomy captures defensive practices against silent failures that crash analysis cannot observe.
Conclusions. Our findings indicate that notebook source should be treated as a confounder in studies of ML developer practice, surface opportunities for notebook tooling, and motivate empirical study of silent ML failures. We release the corpus of 1,092,780 feedback statements and the codebook, to support replication and tooling research.
Cite as
Arumoy Shome, Luís Cruz, Diomidis Spinellis, and Arie van Deursen. Characterizing Feedback Statements in Machine Learning Jupyter Notebooks. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 13:1-13:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{shome_et_al:LIPIcs.ESEM.2026.13,
author = {Shome, Arumoy and Cruz, Lu{\'\i}s and Spinellis, Diomidis and van Deursen, Arie},
title = {{Characterizing Feedback Statements in Machine Learning Jupyter Notebooks}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {13:1--13:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.13},
URN = {urn:nbn:de:0030-drops-279814},
doi = {10.4230/LIPIcs.ESEM.2026.13},
annote = {Keywords: Empirical software engineering, machine learning, Jupyter notebooks, software testing, assertions, mining software repositories}
}
Document
Technical Track Paper
Authors:
Claudia Negri-Ribalta, Ioana Visescu, Muriel-Larissa Frank, Anastasia Sergeeva, Alberto García Simon, Rene Noel, and Lorena Sánchez Chamorro
Abstract
Background. Regulatory data protection requirements (RDPRs) impose requirements on organizations, which translate into technical and organizational measures for their information systems (IS). Meeting these requirements necessitates the collaboration between interdisciplinary teams of lawyers and software engineers. Previous studies indicate that engineers often struggle to interpret and translate RDPRs. Conversely, little is known about the experiences of lawyers when working on RDPRs in interdisciplinary teams, even though their legal interpretations shape the requirements imposed on IS.
Aim. Therefore, this paper seeks to unpack the challenges that lawyers face when collaborating with software engineers to implement data protection requirements and to put them in motion in case of privacy breaches.
Method. Through a deductive qualitative analysis, this paper analyzes semi-structured interviews with 70 data protection lawyers from 25 countries.
Results. Our findings highlight that lawyers throughout regions and regardless of their cultural background describe similar challenges when collaborating with engineers: a lack of shared language and mental models. Lawyers also point out that engineers misunderstand RDPRs for various reasons, not just a lack of education. Furthermore, we identify the key elements that hinder or facilitate collaboration and their implications for privacy engineering. Conclusions: This study is a large qualitative research with participants from multiple countries. Our results suggest that lawyers perceive different challenges when collaborating with software engineers, regardless of their culture. These results suggest these challenges are structural rather than cultural. Organizations should prioritize developing methods and artifacts to overcome these challenges.
Cite as
Claudia Negri-Ribalta, Ioana Visescu, Muriel-Larissa Frank, Anastasia Sergeeva, Alberto García Simon, Rene Noel, and Lorena Sánchez Chamorro. Global Miscommunication in Privacy Requirements Engineering: Legal Experts’ Experiences Collaborating with Software Engineers. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 14:1-14:22, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{negriribalta_et_al:LIPIcs.ESEM.2026.14,
author = {Negri-Ribalta, Claudia and Visescu, Ioana and Frank, Muriel-Larissa and Sergeeva, Anastasia and Simon, Alberto Garc{\'\i}a and Noel, Rene and Chamorro, Lorena S\'{a}nchez},
title = {{Global Miscommunication in Privacy Requirements Engineering: Legal Experts’ Experiences Collaborating with Software Engineers}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {14:1--14:22},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.14},
URN = {urn:nbn:de:0030-drops-279828},
doi = {10.4230/LIPIcs.ESEM.2026.14},
annote = {Keywords: data protection, mental models, privacy, collaboration}
}
Document
Technical Track Paper
Authors:
Lukas Boschanski and Marco Vieira
Abstract
Background. GitHub Actions, the most widely used Continuous Integration and Continuous Deployment (CI/CD) platform, is frequently involved in large-scale software supply chain attacks. While prior work has focused on detecting vulnerable CI/CD pipelines, research has rarely considered preventive perspectives such as analyzing Security Best Practices (SBPs) in CI/CD documentation.
Aims. We investigate the completeness of the official GitHub Actions documentation from a security perspective and identify potential improvements to documenting secure CI/CD practices.
Method. We employ an Large Language Model (LLM)-based documentation mining pipeline to extract security advice items verbatim. We derive actionable CI/CD SBPs from the OWASP Top 10 CI/CD Security Risks framework and conduct a qualitative analysis by mapping the extracted advice items against these best practices to assess which best practices are explicitly covered.
Results. Across 639 documentation pages, we identify 459 security-related advice items, of which 288 are actionable, but only 50% of them map to actionable SBPs. These items primarily address insecure configurations and insufficient credential hygiene. Furthermore, only half of all security advice visual alerts have an adequate warning type.
Conclusions. The GitHub Actions documentation lacks a consistent methodology for incorporating SBPs. Moreover, coverage of established CI/CD security best practices is uneven across OWASP risk categories, with several categories receiving little to no actionable guidance.
Cite as
Lukas Boschanski and Marco Vieira. Evaluating CI/CD Security Best Practices in the GitHub Actions Documentation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 15:1-15:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{boschanski_et_al:LIPIcs.ESEM.2026.15,
author = {Boschanski, Lukas and Vieira, Marco},
title = {{Evaluating CI/CD Security Best Practices in the GitHub Actions Documentation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {15:1--15:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.15},
URN = {urn:nbn:de:0030-drops-279830},
doi = {10.4230/LIPIcs.ESEM.2026.15},
annote = {Keywords: CI/CD security, Software documentation, Security best practices}
}
Document
Technical Track Paper
Authors:
Austen Rainer, Andrew Brown, Rebecca Taylor, and Simon Hettrick
Abstract
Background. Anecdotal evidence suggests that the Fortran ecosystem - the language, tooling, codebases and broad community - continues to make a significant international contribution to science and engineering. Yet the common perception appears to be that Fortran is old and obsolete and, as a result, one should divest from it. There is no empirical evidence - e.g., no contemporary or historical survey - to corroborate or contradict such perception. A consequence is that this perception may be not only mistaken but also damaging to the Fortran ecosystem and to the global contribution that the ecosystem continues to make to society, the economy and environment.
Aims. To gather empirical evidence on the current state of the global Fortran ecosystem, in order to corroborate or challenge perceptions about the language, tooling, codebases and community.
Method. We conduct an international survey of the Fortran community, collecting information from 150 respondents across 25 countries.
Results. We report results about codebases (e.g. code size varies from 1KLOC to 5MLOC), community demographics (e.g., most respondents are male; most respondents hold doctorates), challenges (e.g., code complexity, human resource) and perceptions of Fortran (only 15% consider their codebases to be legacy codebases).
Conclusions. Our results provide the first descriptive benchmark of the Fortran ecosystem which can act as a resource or reference for future research and practice. Also, our results suggest that Fortran may be better understood as a highly-valuable, niche, contemporary ecosystem, one which continues to contribute to world-leading research in many science and engineering disciplines. As an "edge case", it is ideally placed to provide unique and insightful contributions to our understanding of (scientific) software engineering, e.g., technical debt, refactoring, code complexity, and the significance of application domain knowledge. Since all codebases, regardless of language, have the potential to become legacy systems in the future, Fortran is actually at the forefront of confronting long-term challenges and consequences of engineering complex (scientific) software systems.
Cite as
Austen Rainer, Andrew Brown, Rebecca Taylor, and Simon Hettrick. What Do We Really Know About Fortran?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 16:1-16:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{rainer_et_al:LIPIcs.ESEM.2026.16,
author = {Rainer, Austen and Brown, Andrew and Taylor, Rebecca and Hettrick, Simon},
title = {{What Do We Really Know About Fortran?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {16:1--16:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.16},
URN = {urn:nbn:de:0030-drops-279847},
doi = {10.4230/LIPIcs.ESEM.2026.16},
annote = {Keywords: Fortran, demographics, software engineering, research software engineer, computational scientist, software sustainability}
}
Document
Technical Track Paper
Authors:
Sai Mallikarjun Galipelli, Cristian-Alexandru Staicu, and Jibesh Patra
Abstract
Background. Gradual typing, as adopted by languages like TypeScript, aims to combine the flexibility of dynamic typing with the safety guarantees of static type systems. However, the deployed gradual typing systems are inherently unsound, meaning in practice, type inconsistencies are common and can hide or cause subtle bugs. This problem is particularly concerning for widely used libraries, where type annotations serve as both documentation and contracts for users. A value with a wrong type in such libraries may trigger cascading failures in downstream code, including unintended control-flow transfers. Existing work has explored testing-based approaches to detect type mismatches. However, these methods fall short in adversarial settings where malicious inputs are intentionally crafted to exploit type inconsistencies.
Aim. This work introduces TypePatrol, the first adversarial testing framework to uncover security-relevant type mismatches.
Method. TypePatrol systematically mutates values in unit tests and observes the runtime types. Our approach focuses on three classes of security-relevant effects: bypassing input sanitization, invoking incorrect methods, and functions returning wrong types.
Results. We apply TypePatrol to 30 popular JavaScript and TypeScript libraries, and discover hundreds of relevant type inconsistencies. In seven cases, we propose fixes through ten pull requests, the majority of which were accepted, underscoring the real-world relevance of our approach.
Conclusions. Crucially, TypePatrol shows that adversarial testing exposes bugs that slip through conventional testing in even the most popular well-maintained projects. We also discuss cases in which type inconsistencies lead to serious security issues, confirming that these bugs are a serious threat to modern software.
Cite as
Sai Mallikarjun Galipelli, Cristian-Alexandru Staicu, and Jibesh Patra. TypePatrol: Adversarial Testing for Uncovering Security-Relevant Type Inconsistencies in JavaScript Libraries. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 17:1-17:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{galipelli_et_al:LIPIcs.ESEM.2026.17,
author = {Galipelli, Sai Mallikarjun and Staicu, Cristian-Alexandru and Patra, Jibesh},
title = {{TypePatrol: Adversarial Testing for Uncovering Security-Relevant Type Inconsistencies in JavaScript Libraries}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {17:1--17:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.17},
URN = {urn:nbn:de:0030-drops-279852},
doi = {10.4230/LIPIcs.ESEM.2026.17},
annote = {Keywords: Type Inconsistencies, Dynamic Types, JavaScript, TypeScript, Adversarial Testing, Fuzzing, Software Security}
}
Document
Technical Track Paper
Authors:
Hasen Özaytürk and Feza Buzluca
Abstract
Background. Multi-agent approaches to Large Language Model (LLM)-based agentic software engineering, specifically in automated program repair and issue resolution, have converged on architectural patterns that favour isolation: role specialisation, task decomposition in isolated worktrees, or best-of-N with external selection. Collaboration of homogeneous agents on a single task in a shared Git workspace remains understudied, and is actively discouraged in production guidance due to file-level write collisions.
Aims. We investigate whether a shared-workspace configuration can be made viable and effective for autonomous bug fixing and patch production through action-space coordination - a paradigm in which agents observe one another’s commits, lifecycle transitions, and recent test or error outputs rather than exchanging natural-language messages.
Method. We propose PASC (Peer-aware Action-Space Coordination), a layer in which two homogeneous LLM agents share a single Docker container and Git tree. Each agent’s effects are auto-committed under its identity, and each next observation is prepended with a structured peer-activity block. The final submission is the team patch derived from the shared history.
Results. We evaluated PASC on the full Python subset of SWE-Bench Pro using two independently-developed LLMs, against an isolated single-agent baseline and a "silent" two-agent baseline without the peer-activity feed. PASC delivers a statistically significant lift over the single-agent baseline on both models. Crucially, the silent baseline is statistically equivalent to the single-agent one, confirming that the gain stems from action-space coordination rather than parallelism. PASC also outperforms a deployable federated alternative in which agents share the peer-activity feed but work in isolated worktrees. Relative to the silent baseline, PASC reduces cost per resolved task by ∼20% and destructive concurrent edits by ∼47%. Preliminary observations indicate that beyond two agents, active interference increases several-fold, suggesting that larger populations will require supplementary coordination mechanisms.
Conclusions. A shared-workspace configuration of homogeneous agents coordinated through action-space observation improves effectiveness over an isolated single-agent baseline, and does so more economically than a comparable multi-agent baseline without peer awareness.
Cite as
Hasen Özaytürk and Feza Buzluca. Coordinating Agents on a Shared Git Workspace: An Empirical Study of Action-Space Observation for Agentic Software Engineering. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 18:1-18:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{ozayturk_et_al:LIPIcs.ESEM.2026.18,
author = {\"{O}zayt\"{u}rk, Hasen and Buzluca, Feza},
title = {{Coordinating Agents on a Shared Git Workspace: An Empirical Study of Action-Space Observation for Agentic Software Engineering}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {18:1--18:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.18},
URN = {urn:nbn:de:0030-drops-279865},
doi = {10.4230/LIPIcs.ESEM.2026.18},
annote = {Keywords: Multi-agent systems, agentic software engineering, large language model agents, shared workspace coordination, SWE-Bench Pro, empirical software engineering}
}
Document
Technical Track Paper
Authors:
Mussammat Maimuna Faria, Mashiat Amin Farin, Yasin Sazid, and Ahmedul Kabir
Abstract
Background. Generating behavioral models like state transition diagrams from natural language requirements presents challenges in requirements analysis and design. Traditional NLP and ML methods struggle to maintain the correct execution flow and structural consistency. Recent advances in LLMs offer new opportunities, though their effectiveness in generating behavioral models, particularly concerning prompting, retrieval, and repair strategies, remains largely unexamined.
Aims. This paper evaluates how LLM-based generation strategies (zero-shot, one-shot, few-shot, RAG) affect state transition diagram quality (correctness, completeness, understandability, terminological alignment) and whether iterative repair improves validity.
Method. We built a dataset of 80 requirement–diagram pairs from software engineering textbooks. We generated PlantUML code for each requirement using four prompting strategies across four open-source LLMs. To isolate the effect of different knowledge sources, we conducted a RAG analysis study comparing four retrieval corpus configurations (full, examples-only, rules-only, theory-only). We also assessed an iterative repair workflow using automated evaluation. We evaluated syntactic validity via PlantUML validation and structural validity via rule-based analysis, supplementing with human evaluation of quality dimensions. We produced 1,169 PlantUML outputs, of which 120 validated outputs were manually evaluated.
Results. Few-shot prompting improved syntactic validity from 28.7% to 93.5% on average over zero-shot. Iterative repair significantly improved structural validity, especially for DeepSeek (from 22.2% to 88.9%), with most repairs succeeding within two iterations. Human evaluation showed understandability was rated higher than correctness, indicating a quality-perception gap.
Conclusions. We recommend few-shot prompting with iterative repair for AI-assisted behavioral modeling. Our findings reveal critical trade-offs: retrieval improves syntax but risks hallucination, model architecture matters more than strategy, and automated validation cannot substitute for human correctness assessment. This is the first empirical study systematically evaluating prompting, retrieval, and repair in LLM-based UML state transition diagram generation from natural language.
Cite as
Mussammat Maimuna Faria, Mashiat Amin Farin, Yasin Sazid, and Ahmedul Kabir. Towards Reliable AI-Assisted Behavioral Modeling: Evaluating Prompting, Retrieval, and Repair Strategies for UML State Transition Diagram Generation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 19:1-19:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{faria_et_al:LIPIcs.ESEM.2026.19,
author = {Faria, Mussammat Maimuna and Farin, Mashiat Amin and Sazid, Yasin and Kabir, Ahmedul},
title = {{Towards Reliable AI-Assisted Behavioral Modeling: Evaluating Prompting, Retrieval, and Repair Strategies for UML State Transition Diagram Generation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {19:1--19:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.19},
URN = {urn:nbn:de:0030-drops-279873},
doi = {10.4230/LIPIcs.ESEM.2026.19},
annote = {Keywords: State Transition Diagrams, Behavioral Software Modeling}
}
Document
Technical Track Paper
Authors:
Yutong Huo and Dongcheng Li
Abstract
Background. REST APIs have become the primary interaction interface for modern software systems, making automated black-box API testing increasingly important. However, existing testing techniques still face three major limitations: over-reliance on dependencies extracted from static specifications, failure to fully utilise failed requests as feedback, and an overemphasis on crash-oriented verification, resulting in insufficient exploration of non-crash-prone logical defects.
Aims. We design, implement and evaluate a feedback-driven black-box testing framework that turns runtime feedback from a passive outcome into an active control signal, maximising logical-defect yield within practical budgets.
Method. We propose FDRRestTest, which integrates OpenAPI-based static semantic bootstrapping, runtime dependency evolution, a 7-category failure diagnosis and repair loop, a multi-dimensional logical oracle engine and a budget-aware utility scheduler in one closed loop. Under a fixed 600-second time budget, we evaluated FDRRestTest against five state-of-the-art baselines (RESTler, EvoMaster, Morest, ARAT-RL, AutoRestTest) on 12 real-world REST services along five protocols: fixed-time cross-tool comparison, fixed-request cross-tool replication, ablation across three dependency profiles, single-host budget sensitivity, and per-service comparison; significance via two-sided paired Wilcoxon signed-rank tests with Holm-Bonferroni correction, magnitude via standardised paired effect sizes.
Results. FDRRestTest improves fault-revealing efficiency (FRE) by up to 6.3% over the strongest baseline, AutoRestTest (0.9878 vs. 0.9294; an upper estimate, as AutoRestTest’s FRE is a post-hoc lower bound), while issuing 6.3-11.8× fewer requests than AutoRestTest across the 12 services (8.3× on Features Service). The +12.2% logical-yield advantage holds on every one of the 12 services (+7.0%-+16.5%, geometric mean +12.4%), and the FRE gain is consistent across all 12 in per-service means and statistically significant (p_{adj} < .005, d_z = 3.73 on logical bug yield, LBY). An ablation across six subjects spanning three dependency profiles confirms no single component accounts for the advantage, the oracle contributing most - removing it drops FRE below the strongest baseline on all six - while a Pareto-scheduling variant is 4.4% below the weighted utility on the most state-rich host. A manual confirmation study (N = 120, stratified over the 236 deduplicated findings of all six tools, 30 per oracle category) confirms 96/120 (80%) as true defects (Cohen’s κ = 0.78), supporting the oracle-finding interpretation of the LBY gain.
Conclusions. Black-box REST API testing heavily benefits from a tightly coupled closed feedback loop. Within this loop, runtime dependency evolution, failure repair, logical-defect-aware validation, and budget-aware scheduling are each necessary; quantifying their interaction would require multi-component removals, which we leave to future work.
Cite as
Yutong Huo and Dongcheng Li. FDRRestTest: Feedback-Driven Logical Testing for REST APIs. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 20:1-20:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{huo_et_al:LIPIcs.ESEM.2026.20,
author = {Huo, Yutong and Li, Dongcheng},
title = {{FDRRestTest: Feedback-Driven Logical Testing for REST APIs}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {20:1--20:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.20},
URN = {urn:nbn:de:0030-drops-279886},
doi = {10.4230/LIPIcs.ESEM.2026.20},
annote = {Keywords: REST API testing, black-box testing, logical defects, feedback-driven testing, oracle-based validation}
}
Document
Technical Track Paper
Authors:
Shanggui Zhan, Xingqi Wang, Dan Wei, and Xin Xiang
Abstract
Background. LLM-as-Judge is increasingly adopted in automated program repair (APR) to assess patch correctness, yet its validity as a measurement instrument remains unaudited: potential benchmark-identifier leakage, evidence-presentation bias, and explanation unreliability have not been systematically examined.
Aims. We investigate whether LLM Judge produces unbiased, stable, and well-grounded patch correctness judgments across three validity dimensions: metadata leakage, evidence presentation, and explanation reliability.
Method. On a balanced, paired dataset of 326 patches from 163 Defects4J bugs, we evaluate GPT-4o and DeepSeek-V3 across a 12-setting prompt matrix covering benchmark-identifying metadata exposure and evidence presentation variants. We further conduct a manual analysis of 60 GPT-4o explanations using a five-category failure taxonomy and three independent annotators.
Results. Under sanitized conditions, GPT-4o achieves an accuracy of 77.3% (MCC = 0.575, FPR = 17.2%), while DeepSeek-V3 achieves 80.4% accuracy (MCC = 0.625, FPR = 22.1%), with no statistically significant difference between the two models. Exposing combined benchmark-identifying metadata yields a suggestive accuracy increase of 4.3 percentage points for GPT-4o and 1.8 percentage points for DeepSeek-V3. Declaring that all tests have passed does not significantly bias either model; in contrast, providing a developer reference patch leads to substantial and statistically significant improvements for both models. Explanation failures are concentrated in misclassified cases, with the false-positive quadrant showing a 93.3% failure rate, primarily driven by hallucinated evidence and missed edge cases.
Conclusions. LLM Judge is useful as a triage aid but is insufficient as a standalone correctness oracle. Reliable deployment requires prompt sanitization, explicit false positive rate reporting, clear distinction between reference-assisted and standalone assessment, and treating LLM explanations as investigative hypotheses rather than self-validating justifications.
Cite as
Shanggui Zhan, Xingqi Wang, Dan Wei, and Xin Xiang. How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 21:1-21:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{zhan_et_al:LIPIcs.ESEM.2026.21,
author = {Zhan, Shanggui and Wang, Xingqi and Wei, Dan and Xiang, Xin},
title = {{How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {21:1--21:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.21},
URN = {urn:nbn:de:0030-drops-279898},
doi = {10.4230/LIPIcs.ESEM.2026.21},
annote = {Keywords: automated program repair, patch correctness assessment, LLM-as-Judge, empirical study}
}
Document
Technical Track Paper
Authors:
Enrique Barba Roque, Luís Cruz, and Annibale Panichella
Abstract
Background. Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network.
Aims. This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric.
Method. We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization.
Results. Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90% and memory usage by 86%, with only modest accuracy trade-offs.
Conclusions. Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.
Cite as
Enrique Barba Roque, Luís Cruz, and Annibale Panichella. Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 22:1-22:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{barbaroque_et_al:LIPIcs.ESEM.2026.22,
author = {Barba Roque, Enrique and Cruz, Lu{\'\i}s and Panichella, Annibale},
title = {{Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {22:1--22:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.22},
URN = {urn:nbn:de:0030-drops-279905},
doi = {10.4230/LIPIcs.ESEM.2026.22},
annote = {Keywords: Knowledge distillation, Green AI, Many-objective Optimization, LLMs for Code, FLOPs, AI for SE}
}
Document
Technical Track Paper
Authors:
Vitor Gaboardi dos Santos, Boualem Benatallah, and Silvana Togneri MacMahon
Abstract
Background. Users can express the same request in many ways when interacting with LLM agents, particularly non-native English speakers whose writing reflects linguistic patterns transferred from their first language (L1). Existing benchmarks for evaluating API usage in LLM-based software systems typically rely on standardized English and overlook linguistic diversity. At the same time, collecting such data from users with diverse linguistic backgrounds is costly and difficult to scale.
Aims. This study proposes L1-AUG, an L1-aware utterance generation method, and investigates three research questions: whether LLMs can generate API-calling utterances reproducing linguistic patterns of non-native English speakers; whether L1-aware utterance generation leads to higher linguistic diversity; and whether API-calling utterances that simulate non-native English linguistic patterns are more challenging for LLM agents to solve.
Results. L1-AUG can generate utterances reflecting linguistic patterns extracted from essays written by non-native English speakers; L1-AUG generates benchmarks with higher linguistic diversity; and LLM agents achieve lower performance on the benchmark created using L1-AUG.
Conclusions. Conditioning benchmark generation on linguistic patterns extracted from non-native English essays produces benchmarks that better reflect users with diverse linguistic backgrounds. Our results indicate that LLM agents struggle more with such inputs, suggesting that training on linguistically diverse datasets may be necessary to improve robustness in API-calling tasks.
Cite as
Vitor Gaboardi dos Santos, Boualem Benatallah, and Silvana Togneri MacMahon. Beyond Standard English: L1-Aware Benchmarks for Evaluating API-Calling in LLM Agents. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 23:1-23:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{gaboardidossantos_et_al:LIPIcs.ESEM.2026.23,
author = {Gaboardi dos Santos, Vitor and Benatallah, Boualem and MacMahon, Silvana Togneri},
title = {{Beyond Standard English: L1-Aware Benchmarks for Evaluating API-Calling in LLM Agents}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {23:1--23:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.23},
URN = {urn:nbn:de:0030-drops-279916},
doi = {10.4230/LIPIcs.ESEM.2026.23},
annote = {Keywords: Software Testing Benchmarks, LLM-based Software Agents, Linguistic Diversity, L1 Transfer Patterns}
}
Document
Technical Track Paper
Authors:
Ayane Shirakawa, Tatsuya Shirai, Yutaro Kashiwa, Masanari Kondo, Yasutaka Kamei, and Hajimu Iida
Abstract
Background. Continuous Integration (CI) aims to shorten release cycles by automating tests on every change. When the same test method keeps failing across consecutive revisions, new defects get buried among existing failures, dulling developer vigilance and weakening CI’s benefit of rapid bug localization.
Aims. We define Test Alert Snooze as the state in which the same test method fails across two or more consecutive revisions, and provide its first empirical characterization: its prevalence, persistence in revisions and elapsed time, the commits developers make during Test Alert Snooze, and what they change to resolve it.
Method. We analyzed 27 open-source Python and Java projects on GitHub Actions with high test execution frequencies. From build histories and logs over a 90-day window (December 21, 2025 to March 20, 2026), we identified test methods that failed across two or more consecutive revisions, and classified both the commits made during Test Alert Snooze and those that resolved it.
Results. Test Alert Snooze accounts for 42.9% of observed test failures, with a median persistence of 2 consecutive revisions and about one day; extreme cases reached 11 commits and 16 days. During Test Alert Snooze, fix commits made up only 8.2% of developer activity, while docs, test, refactor, and feat collectively dominated. Resolutions were mostly single-category modifications to Test or Product files, and among resolving commits test was the most frequent (32.4%), well above fix (11.8%).
Conclusions. Test-failure persistence is not a single phenomenon but differs across the build, job, and test-method levels. At the test-method level, resolving a failure often takes the form of test-code maintenance rather than bug fixing, so predicting, detecting, and repairing Test Alert Snooze calls for techniques targeting test code alongside product code. This first characterization lays the groundwork for such techniques and for CI features that visualize test-method-level failure persistence.
Cite as
Ayane Shirakawa, Tatsuya Shirai, Yutaro Kashiwa, Masanari Kondo, Yasutaka Kamei, and Hajimu Iida. Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 24:1-24:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{shirakawa_et_al:LIPIcs.ESEM.2026.24,
author = {Shirakawa, Ayane and Shirai, Tatsuya and Kashiwa, Yutaro and Kondo, Masanari and Kamei, Yasutaka and Iida, Hajimu},
title = {{Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {24:1--24:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.24},
URN = {urn:nbn:de:0030-drops-279929},
doi = {10.4230/LIPIcs.ESEM.2026.24},
annote = {Keywords: Continuous integration (CI), Test failures, Empirical analysis}
}
Document
Technical Track Paper
Authors:
Quanhe Wang, Cheng Wen, Dugang Liu, Xingjian Han, Bin Yu, Ping Chen, Shengchao Qin, and Cong Tian
Abstract
Background. Large language models (LLMs) can generate executable code, but their reliability across programming languages remains difficult to characterize.
Aims. We investigate multilingual code capability under problem alignment, focusing on functional reliability, cross-language consistency, translation behavior, problem and sampling effects, and execution quality.
Method. We present MciBench, covering 1,489 programming problems and eight languages, of which 1,466 have execution-validated coverage in all eight languages. We evaluate representative LLMs through multilingual generation and analyze source-code-conditioned translation on 100 randomly selected all-language-covered problems.
Results. Performance varies substantially across models and target languages, while source-code-conditioned translation also exhibits marked variation across source-target settings. Problem difficulty and candidate budget materially affect correctness, strong aggregate performance does not necessarily imply cross-language consistency, and accepted programs differ in runtime and memory efficiency.
Conclusions. Aggregate pass@k alone is insufficient for characterizing multilingual code capability, motivating evaluation that considers language consistency, problem sensitivity, sampling behavior, translation settings, and post-correctness execution quality.
Cite as
Quanhe Wang, Cheng Wen, Dugang Liu, Xingjian Han, Bin Yu, Ping Chen, Shengchao Qin, and Cong Tian. An Empirical Study of Problem-Aligned Multilingual Code Generation and Directed Code Translation by Large Language Models. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 25:1-25:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{wang_et_al:LIPIcs.ESEM.2026.25,
author = {Wang, Quanhe and Wen, Cheng and Liu, Dugang and Han, Xingjian and Yu, Bin and Chen, Ping and Qin, Shengchao and Tian, Cong},
title = {{An Empirical Study of Problem-Aligned Multilingual Code Generation and Directed Code Translation by Large Language Models}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {25:1--25:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.25},
URN = {urn:nbn:de:0030-drops-279936},
doi = {10.4230/LIPIcs.ESEM.2026.25},
annote = {Keywords: Large language models, code generation, code translation, multilingual programming, empirical software engineering}
}
Document
Technical Track Paper
Authors:
Ahmed Belhouchette, Moataz Chouchen, Marouene Chaieb, Mohammad Hamdaqa, and Abdelwahab Hamou-Lhadj
Abstract
Background. Developers increasingly coordinate dependent review workflows by submitting sequences of related changes rather than monolithic ones. In Gerrit, these dependencies form relation chains: structured review units that link changes together. As chains become more common, they shape review activities through synchronization overhead, CI amplification, and merge-ordering constraints.
Aim. We investigate how developers adopt relation chains and how these dependency structures influence review dynamics and outcomes.
Method. We analyze 29,580 relation chains from 15 repositories across three Gerrit ecosystems (OpenStack, Wikimedia, ONAP), comprising 401,256 changes, using Mann-Kendall trend tests, Mann-Whitney with Cliff’s δ for chain-vs-solo comparison, and Spearman correlations for base-descendant dependency.
Results. Chain prevalence ranges from 5% to 49% across projects, increasing in 14 of 15. Chain changes take a median of 2.6× longer to merge than size-matched solo changes, with the gap widening for very large changes. Review effort propagates through dependency-linked review workflows: base-change review activity co-varies with descendant review activity (ρ = 0.43-0.61 in 14-15/15 projects), and 33.5% of chain members undergo structural evolution during review.
Conclusions. Relation chains operate as durable, ecosystem-shaped coordination units with internal structure that change-centric analyses cannot capture. Future review analytics, reviewer-assignment systems, and AI-assisted review tools should reason over chains rather than isolated changes.
Cite as
Ahmed Belhouchette, Moataz Chouchen, Marouene Chaieb, Mohammad Hamdaqa, and Abdelwahab Hamou-Lhadj. How Developers Use Relation Chains in Gerrit-Based Review Ecosystems: An Empirical Study Across Three Open-Source Ecosystems. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 26:1-26:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{belhouchette_et_al:LIPIcs.ESEM.2026.26,
author = {Belhouchette, Ahmed and Chouchen, Moataz and Chaieb, Marouene and Hamdaqa, Mohammad and Hamou-Lhadj, Abdelwahab},
title = {{How Developers Use Relation Chains in Gerrit-Based Review Ecosystems: An Empirical Study Across Three Open-Source Ecosystems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {26:1--26:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.26},
URN = {urn:nbn:de:0030-drops-279942},
doi = {10.4230/LIPIcs.ESEM.2026.26},
annote = {Keywords: Code review, relation chains, stacked changes, dependent patches, empirical software engineering, Gerrit}
}
Document
Technical Track Paper
Authors:
Keerthiga Rajenthiram, Faezeh Amou Najafabadi, Ilias Gerostathopoulos, and Patricia Lago
Abstract
Background. Automated Machine Learning (AutoML) is increasingly embedded in software and data engineering workflows, automating the design and optimization of machine learning (ML) pipelines. AutoML has been extensively synthesized through secondary studies that map tools, algorithms, and evaluation practices. However, these reviews predominantly reflect a research-centric perspective and provide limited evidence on how AutoML is actually used, adapted, and experienced in practice. Traditional surveys and interviews capture only a narrow slice of the practitioner community, leaving open how well academic claims about AutoML align with day-to-day engineering realities.
Aims. This study aims to (i) assess to what extent AutoML practices and challenges reported in existing secondary studies reflect practitioners' experiences, and (ii) identify mismatches between research and practice by empirically comparing academic claims with practitioner experiences shared on YouTube.
Method. Following established guidelines for multivocal literature reviews in software engineering, we conduct a multivocal evidence synthesis combining 16 peer-reviewed secondary studies (systematic and multivocal literature reviews) with 30 practitioner-created YouTube videos in which engineers, data scientists, and ML practitioners describe their AutoML workflows, workarounds, and pain points.
Results. Our analysis reveals a marked imbalance: academic reviews emphasize algorithm selection, search strategies, and benchmarking, whereas practitioners foreground data quality and cleaning, resource and cost constraints, tool robustness, scalability, and integration into production pipelines. We also observe areas of alignment, for example around hyperparameter optimization and the need for reproducible workflows.
Conclusions. Treating YouTube videos as curated gray literature under established multivocal review guidelines, our study shows that they provide complementary, practice-grounded evidence about AutoML adoption. The identified gaps and alignments suggest concrete directions for empirical AutoML research that more directly address operational and organizational constraints in real-world software engineering settings.
Cite as
Keerthiga Rajenthiram, Faezeh Amou Najafabadi, Ilias Gerostathopoulos, and Patricia Lago. A Multivocal Review of AutoML Practices, Challenges, Opportunities and Open Issues. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 27:1-27:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{rajenthiram_et_al:LIPIcs.ESEM.2026.27,
author = {Rajenthiram, Keerthiga and Najafabadi, Faezeh Amou and Gerostathopoulos, Ilias and Lago, Patricia},
title = {{A Multivocal Review of AutoML Practices, Challenges, Opportunities and Open Issues}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {27:1--27:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.27},
URN = {urn:nbn:de:0030-drops-279953},
doi = {10.4230/LIPIcs.ESEM.2026.27},
annote = {Keywords: Automated Machine Learning (AutoML), Multivocal Literature Review, YouTube Videos, Gray Literature, Empirical Study, Research–Practice Gap}
}
Document
Technical Track Paper
Authors:
Ziyu Chen, Qiangpu Chen, Yuwei Li, Taiyan Wang, Shiwen Ou, Qingsong Xie, Lu Zhang, and Zulie Pan
Abstract
Background. Agents have recently improved LLM-based vulnerability detection through retrieval, iterative analysis, and verifier modules. However, existing evaluations mainly report improvements in detection performance without determining whether these improvements reflect more reliable vulnerability reasoning or access to retrieved context.
Aims. This study investigates how different agent components affect the reasoning-oriented evaluation dimensions of agent-based vulnerability detection.
Method. We present a controlled measurement framework with six conditions and evaluate it on six base models. The conditions progress from plain single-pass prompting through matched-context replay and ReAct-style tool use to self-review, trace-isolated verification, and multi-role decomposition, allowing us to separate context effects, iterative orchestration, and added agent modules under a shared protocol. Beyond vulnerability detection performance metrics, our framework evaluates three reasoning-oriented dimensions: robustness under semantics-preserving perturbations, resistance to misleading non-causal context, and grounding in concrete code evidence.
Results. The full agent often achieves higher accuracy or F1 than plain single-pass prompting, but these improvements do not consistently extend to the three measured behavioral dimensions of robustness, resistance, and grounding. The models are not consistently more stable under harmless code changes, are often misled by non-causal prompt cues, and rarely cite the exact code lines that explain the vulnerability. In the strongest improvement case, much of the accuracy change already appears when the model receives the same retrieved tool outputs without the full agent process; after that context is held fixed, the remaining agent effect is mainly associated with predicting vulnerable more often.
Conclusions. Higher detection performance should not be taken as direct evidence of better vulnerability reasoning. In our experiments, better performance metrics can reflect retrieved context or changes in how often the model predicts vulnerable, rather than a clear improvement in stable, evidence-based vulnerability reasoning. Future evaluations should report these behavior checks alongside performance metrics.
Cite as
Ziyu Chen, Qiangpu Chen, Yuwei Li, Taiyan Wang, Shiwen Ou, Qingsong Xie, Lu Zhang, and Zulie Pan. Do LLM Vulnerability-Detection Agents Reason Better? An Empirical Decomposition of Robustness, Resistance, and Grounding. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 28:1-28:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{chen_et_al:LIPIcs.ESEM.2026.28,
author = {Chen, Ziyu and Chen, Qiangpu and Li, Yuwei and Wang, Taiyan and Ou, Shiwen and Xie, Qingsong and Zhang, Lu and Pan, Zulie},
title = {{Do LLM Vulnerability-Detection Agents Reason Better? An Empirical Decomposition of Robustness, Resistance, and Grounding}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {28:1--28:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.28},
URN = {urn:nbn:de:0030-drops-279961},
doi = {10.4230/LIPIcs.ESEM.2026.28},
annote = {Keywords: Large language models, vulnerability detection, software security, agent-based systems, empirical evaluation}
}
Document
Technical Track Paper
Authors:
Feray Gulu-zada and Hina Anwar
Abstract
Background. LLMs are increasingly used for code refactoring, but their effects on multiple software quality dimensions remain unclear. Green coding targets energy efficiency, while clean coding emphasizes readability and maintainability. These objectives may interact or conflict, making LLM-based refactoring a multi-objective problem.
Aims. This study investigates whether prompting can guide LLM-based refactoring toward balanced green and clean code outcomes. We examine how different prompting strategies affect energy consumption, runtime performance, and maintainability.
Method. We conduct a controlled empirical study on 15 human-written Python scripts using four LLMs and five prompting strategies: clean, green, multi-objective, clean-to-green staged, and green-to-clean staged prompting. Generated refactorings are validated for correctness before analysis. Valid outputs are evaluated using CPU energy consumption, runtime, maintainability index, cyclomatic complexity, source lines of code, and Halstead volume. We analyze results using descriptive statistics, statistical tests, correlation analysis, bootstrap confidence intervals, and Pareto-efficiency analysis.
Results. Of 300 generated refactorings, 267 passed validation. Among valid outputs, 53.18% (142 of 267) reduced energy consumption, but the overall energy effect was small, variable, and not statistically significant across prompting strategies. Energy and runtime changes showed a moderate positive correlation, while maintainability showed weaker relationships with efficiency metrics. Prompting strategies significantly affected maintainability but not energy or runtime. Clean and multi-objective prompting produced the most balanced Pareto outcomes, though these differences were mainly driven by maintainability.
Conclusions. LLM-based refactoring can sometimes reduce energy consumption without necessarily degrading maintainability, but its effects are inconsistent and depend on model choice, prompt design, and correctness validation. Prompting is useful for steering LLM-based refactoring, but current evidence supports cautious, validated use rather than assuming reliable multi-objective optimization.
Cite as
Feray Gulu-zada and Hina Anwar. Balancing Green and Clean Code: Prompting for LLM-Based Refactoring. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 29:1-29:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{guluzada_et_al:LIPIcs.ESEM.2026.29,
author = {Gulu-zada, Feray and Anwar, Hina},
title = {{Balancing Green and Clean Code: Prompting for LLM-Based Refactoring}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {29:1--29:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.29},
URN = {urn:nbn:de:0030-drops-279971},
doi = {10.4230/LIPIcs.ESEM.2026.29},
annote = {Keywords: Large language models, code refactoring, energy efficiency, maintainability, green software, prompt engineering}
}
Document
Technical Track Paper
Authors:
Alessandro Aneggi, Vincenzo Stoico, and Andrea Janes
Abstract
Background. Performance antipatterns are well known for degrading the responsiveness of microservice-based systems. However, although their impact on latency has been widely studied, their relationship with energy consumption remains largely unexplored.
Aims. This study empirically investigates whether ten widely recognized performance antipatterns, as defined by Smith and Williams, affect power consumption and energy-related behavior in containerized environments.
Method. We implemented ten antipatterns as isolated containerized services and evaluated them under controlled load conditions. For each antipattern, we conducted 30 repeated runs and collected measurements of response time, CPU and DRAM power consumption, and resource utilization.
Results. While all injected antipatterns degraded performance as expected, only a subset showed a statistically significant positive relationship between increased response time and higher power consumption. Some antipatterns exhibited a clear association between performance degradation and increased power demand. In contrast, others reached CPU saturation early, such that power draw remained nearly constant despite rising response times. In these cases, additional slowdowns primarily increased total execution time rather than instantaneous power demand.
Conclusions. Performance antipatterns do not affect power and energy-related behavior uniformly. This study provides a systematic foundation for characterizing how different performance flaws are associated with power and energy-related behavior, offering actionable insights for architects aiming to design more energy-aware software systems.
Cite as
Alessandro Aneggi, Vincenzo Stoico, and Andrea Janes. Beyond Response Time: Investigating the Power Behavior of Performance Antipatterns. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 30:1-30:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{aneggi_et_al:LIPIcs.ESEM.2026.30,
author = {Aneggi, Alessandro and Stoico, Vincenzo and Janes, Andrea},
title = {{Beyond Response Time: Investigating the Power Behavior of Performance Antipatterns}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {30:1--30:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.30},
URN = {urn:nbn:de:0030-drops-279987},
doi = {10.4230/LIPIcs.ESEM.2026.30},
annote = {Keywords: performance antipatterns, microservices, power consumption, energy consumption, energy efficiency}
}
Document
Technical Track Paper
Authors:
Isabella Graßl
Abstract
Background. Software engineering is a sociotechnical discipline in which teamwork is central, even as AI tools are integrated into professional and educational practice. While prior research has primarily examined AI as an individual productivity tool, we know little about how software teams position AI within collaboration.
Aims. Drawing on boundary work theory, we explore how students integrate AI into project-based software engineering teamwork and how they construct boundaries between AI and human collaboration.
Method. We conducted a qualitative multiple-case study across seven software engineering courses, collecting interviews (n = 27) and reflections (n = 89) from 116 students.
Results. Students construct explicit human–AI boundaries within software engineering teamwork. They draw complementary boundaries, using AI to support individual preparation and thereby improving the conditions around collaboration, and competitive boundaries, actively excluding AI from interpersonal aspects of teamwork that they reserve for humans, grounded in concerns about accountability, trust, and relational value. We found no evidence in students' accounts that they perceived AI as a co-evolving collaborative partner.
Conclusions. We argue that students actively construct AI as infrastructural support rather than as a social collaborator, thereby preserving human responsibility, empathy, and trust within software engineering teamwork. These findings contribute to emerging discussions on human–AI collaboration and the evolving role of software engineers in this era.
Cite as
Isabella Graßl. How Students Construct Human–AI Boundaries in Software Engineering Teamwork. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 31:1-31:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{gral:LIPIcs.ESEM.2026.31,
author = {Gra{\ss}l, Isabella},
title = {{How Students Construct Human–AI Boundaries in Software Engineering Teamwork}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {31:1--31:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.31},
URN = {urn:nbn:de:0030-drops-279994},
doi = {10.4230/LIPIcs.ESEM.2026.31},
annote = {Keywords: Software Engineering, AI, Teamwork, Collaboration}
}
Document
Technical Track Paper
Authors:
Philipp Holzinger and Philipp Rajwa
Abstract
Background. Many computer science papers rely on software prototypes that implement novel concepts or generate data for empirical evaluation. To support confirmatory research and reuse, more venues now request code and data submissions and conduct artifact evaluations. However, concerns about a reproducibility crisis remain, as studies show that only a small share of published artifacts is actually accessible and working for reuse or reproducing results.
Aims. To shed more light on this in the context of Android-related research, we investigate to what extent research artifacts are available, functional, and whether they allow us to reproduce published experimental results. Moreover, we want to understand the reasons why research artifacts might fail to fulfill these objectives.
Method. We conducted a large-scale empirical study of Android-related research artifacts that were published between 2011 and 2025. For an initial set of 269 papers, we attempted to locate their associated artifacts, assess whether they can be obtained and executed, and evaluate whether they reproduce the expected output.
Results. Starting from 269 papers, only 163 artifacts were still available. Among these, 118 papers published experimental results, yet we were only able to establish 10 experimental setups. Ultimately, only 8 artifacts could fully or partially reproduce the expected output.
Conclusions. Our findings provide further evidence for reproducibility challenges in Android-related research. Based on these sobering results, we propose a set of recommendations to strengthen the perceived value of artifacts and to improve their availability, quality, and visibility.
Cite as
Philipp Holzinger and Philipp Rajwa. Large-Scale Evaluation of Android-Related Research Artifacts. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 32:1-32:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{holzinger_et_al:LIPIcs.ESEM.2026.32,
author = {Holzinger, Philipp and Rajwa, Philipp},
title = {{Large-Scale Evaluation of Android-Related Research Artifacts}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {32:1--32:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.32},
URN = {urn:nbn:de:0030-drops-280000},
doi = {10.4230/LIPIcs.ESEM.2026.32},
annote = {Keywords: research artifacts, reproducibility, reuse, Android, empirical study}
}
Document
Technical Track Paper
Authors:
Xing Cui, Jingzheng Wu, Tianyue Luo, and Xiang Ling
Abstract
Background. Agent systems increasingly invoke heterogeneous third-party services through the Model Context Protocol (MCP), where service licenses, model licenses, and data licenses may be jointly applied during runtime execution. This emerging execution paradigm introduces cross-service compliance conflicts that are difficult to capture with conventional license analysis methods. Existing methods usually assume that the license set under analysis is statically predetermined, making them insufficient for handling novel term semantics in MCP licenses and for analyzing conflict patterns that arise from task-specific selected service subsets.
Aims. This paper investigates compositional license conflicts in MCP-based agent workflows. We seek to characterize how license terms interact across compositional MCP services and to provide an automated analysis framework that detects conflicts and identifies license-compatible service replacements.
Method. We propose MCP-Licentra, a license conflict analysis framework for MCP-based agent workflows. MCP-Licentra contains four integrated components. First, it extends an existing license term taxonomy with MCP-specific categories and constructs annotated training data for domain adaptation. Second, it fine-tunes License-Llama3-8B with supervised fine-tuning (SFT) and odds ratio preference optimization (ORPO) to recognize license terms and their associated attitudes. Third, it formulates task-specific service selection as a power-set legality verification problem, enabling conflict checking across possible service invocation subsets. Fourth, it recommends license-compatible replacements for conflicting services through semantic embedding retrieval and an iterative compliance checking algorithm under joint functional similarity and license compatibility constraints.
Results. We evaluate MCP-Licentra on 3,811 real-world MCP services and 500 compositional service scenarios. MCP-Licentra achieves F1-scores of 88.82% and 86.44% for term identification and license understanding, respectively. For group-level conflict detection, it obtains an FPR of 3.15% and an FNR of 8.77%. For compositional conflict detection, it achieves an overall FPR of 4.90% and an FNR of 3.33%. In addition, MCP-Licentra reaches a resolution success rate of 71.23%, outperforming all baselines.
Conclusion. The results show that MCP-based agent workflows introduce compositional compliance risks that are difficult to capture with conventional static license analysis. MCP-Licentra provides an automated framework and empirical evidence for detecting and mitigating such risks in agent systems.
Cite as
Xing Cui, Jingzheng Wu, Tianyue Luo, and Xiang Ling. Blurred Compliance Boundaries: Analyzing Open Source License Conflicts in MCP-Based Agent Workflows. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 33:1-33:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{cui_et_al:LIPIcs.ESEM.2026.33,
author = {Cui, Xing and Wu, Jingzheng and Luo, Tianyue and Ling, Xiang},
title = {{Blurred Compliance Boundaries: Analyzing Open Source License Conflicts in MCP-Based Agent Workflows}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {33:1--33:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.33},
URN = {urn:nbn:de:0030-drops-280017},
doi = {10.4230/LIPIcs.ESEM.2026.33},
annote = {Keywords: MCP, license conflict detection, compositional license analysis, large language model, agent compliance}
}
Document
Technical Track Paper
Authors:
Syed Fakhar Abbas Naqvi and Hina Anwar
Abstract
Background. Large Language Models (LLMs) are increasingly used in software engineering tasks such as code generation, bug detection, and program repair. However, their use introduces inference-time costs that are rarely considered together with functional performance, especially in debugging and repair workflows where models may be invoked repeatedly.
Aim. This study evaluates the correctness-energy trade-offs of LLMs across model scales in Python bug-fixing tasks. We examine how model size relates to code correctness, inference time, and GPU energy consumption.
Method. We conduct a controlled empirical evaluation of six open-source code-oriented LLMs, ranging from 1.5B to 15B parameters, on 40 real-world Python bugs from the BugsInPy benchmark. Each model is executed multiple times per task under a consistent hardware and prompting setup. Repair performance is measured using Pass@k and test-suite outcomes, while computational cost is measured using inference time and GPU energy consumption. We also analyze generated outputs to identify recurring failure patterns.
Results. Increasing model size substantially raised computational cost without consistently improving correctness. The smallest model, Qwen 1.5B, achieved the best overall performance, with a Pass@1 score of 43% and an average runtime of 38.7 seconds. In contrast, StarCoder 15B achieved a Pass@1 score of 23%, required 96.7 seconds, and consumed approximately 7 times more energy per inference while solving fewer tasks. Higher success rates were also observed in web-related projects, suggesting that structurally localized bugs are more effectively handled by current models.
Conclusion. On the evaluated localized Python bug-fixing tasks, smaller LLMs offered a more favorable trade-off between correctness, inference time, and energy consumption. These results suggest that scaling model size does not automatically lead to more efficient LLM-based repair, and that inference cost should be considered alongside functional correctness when comparing models.
Cite as
Syed Fakhar Abbas Naqvi and Hina Anwar. Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 34:1-34:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{naqvi_et_al:LIPIcs.ESEM.2026.34,
author = {Naqvi, Syed Fakhar Abbas and Anwar, Hina},
title = {{Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {34:1--34:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.34},
URN = {urn:nbn:de:0030-drops-280023},
doi = {10.4230/LIPIcs.ESEM.2026.34},
annote = {Keywords: Large Language Models, Bug Fixing, Energy Consumption, Empirical Software Engineering, Sustainability}
}
Document
Technical Track Paper
Authors:
Ragib Shahariar Ayon, Rayed Fahmi, Sumon Biswas, and Shibbir Ahmed
Abstract
Background. Evaluating automated Software Requirements Specification (SRS) generation is challenging because few datasets provide fine-grained traceability between source requirements, intermediate elicitation artifacts, and generated specifications.
Aims. We study whether legacy SRS documents can serve as controlled reference evidence for constructing traceable synthetic pre-SRS artifacts that enable fine-grained evaluation of LLM-based requirements transformations.
Method. We conduct an empirical study using ReqGenX, a controlled pipeline that decomposes SRS sections into atomic statements grounded in the SRS content, routes atoms to standards-inspired artifact types through multi-LLM plurality voting, and generates artifacts using constrained prompts with iterative judge-guided refinement. We evaluate ReqGenX on seven PURE SRS documents using grounding, quality, information retention, and downstream reconstruction analyses.
Results. ReqGenX produces highly grounded, well-formed atoms that score strongly on our atomicity and downstream-utility criteria, with median AlignScore values typically between 0.96 and 0.99 and Prometheus scores ranging from 4.34 to 4.85. Generated artifacts remain strongly grounded in their source atoms, with AlignScore values typically between 0.80 and 0.94 and judge pass rates near 100%; stricter Prometheus evaluation yields pass rates ranging from 54.8% to 97.1%. In a downstream SRS reconstruction case study, artifact-backed atoms remain recoverable from generated SRSs, with SBERT means between 0.69 and 0.75 and AlignScore medians between 0.76 and 0.84.
Conclusions. Traceable synthetic pre-SRS artifacts can support fine-grained evaluation of LLM-based SRS generation, while exposing tradeoffs among faithfulness, information retention, and artifact completeness. These artifacts are intended as controlled evaluation resources rather than substitutes for primary elicitation evidence.
Cite as
Ragib Shahariar Ayon, Rayed Fahmi, Sumon Biswas, and Shibbir Ahmed. ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 35:1-35:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{ayon_et_al:LIPIcs.ESEM.2026.35,
author = {Ayon, Ragib Shahariar and Fahmi, Rayed and Biswas, Sumon and Ahmed, Shibbir},
title = {{ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {35:1--35:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.35},
URN = {urn:nbn:de:0030-drops-280036},
doi = {10.4230/LIPIcs.ESEM.2026.35},
annote = {Keywords: Requirements, large language model, software engineering}
}
Document
Technical Track Paper
Authors:
Suzuka Yoshimoto, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, and Hajimu Iida
Abstract
Background. When developers write a TODO or FIXME comment, they are explicitly admitting that the code is suboptimal: a built-in warning that this logic deserves extra scrutiny. Yet it is an open question whether Self-Admitted Technical Debt (SATD) actually receives that scrutiny in the form of software testing.
Aim. We aim to characterize the relationship between SATD and testing across three dimensions: the extent to which SATD-affected code is covered by existing tests, whether developers synchronize test additions with debt resolution, and whether such testing affects the long-term observability of resulting defects.
Method. For that, we conducted an empirical study on eight open-source Java projects, analyzing test coverage of 784 SATD instances identified in the latest releases and performing a longitudinal examination of 5,175 SATD removal events.
Results. Our results show that while 60.7% of SATD-affected code is covered by existing test suites, developers rarely synchronize test modifications with debt resolution; manual inspection confirms that only 3.4% of SATD removal commits include new tests specifically targeting the resolved debt (vs. 12.5% that co-add tests in the same commit). Longitudinal analysis further suggests that SATD resolutions exhibit nearly identical localized bug induction rates within short-to-medium-term windows regardless of test modifications. However, over a longer, unrestricted observation window, a slight divergence emerges where the test-added group reaches a higher cumulative defect alignment probability (6.32% vs. 4.37%), a counterintuitive trend potentially driven by the selective testing of inherently complex components.
Conclusion. Developers treat SATD repayment as an ordinary code change rather than as a high-risk maintenance activity: most debt removals proceed without targeted verification, despite the developer’s own prior flag that the code is suboptimal.
Cite as
Suzuka Yoshimoto, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, and Hajimu Iida. Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-Change, and Impact. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 36:1-36:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{yoshimoto_et_al:LIPIcs.ESEM.2026.36,
author = {Yoshimoto, Suzuka and Horikawa, Kosei and Feitosa, Daniel and Kashiwa, Yutaro and Iida, Hajimu},
title = {{Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-Change, and Impact}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {36:1--36:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.36},
URN = {urn:nbn:de:0030-drops-280048},
doi = {10.4230/LIPIcs.ESEM.2026.36},
annote = {Keywords: Self-Admitted Technical Debt, Test Coverage, Software Testing}
}
Document
Technical Track Paper
Authors:
Alexander Puma Pucho and Bruno B. P. Cafeo
Abstract
Background. LLM-based multi-agent systems (LLM-MAS) are increasingly applied to software engineering (SE) tasks. However, most systems are evaluated end-to-end, entangling architectural choices with model selection, prompting, and task-specific tooling. As a consequence, the effects of individual architectural factors remain empirically underexplored.
Aims. We investigate how Communication Structure (Hierarchical vs. Centralized) and Coordination Strategy (Dynamic vs. Static) influence three outcome dimensions in LLM-MAS for Python refactoring: output generation and validity; operational cost and execution efficiency; and structural transformation. We also examine whether the two factors interact.
Method. We conduct a controlled 2 × 2 factorial experiment under identical model, dataset, orchestration, prompting, and tool configurations, evaluated on 2,000 real-world Python refactoring instances drawn from open-source machine learning (ML) repositories. We measure output generation and validity through file generation, syntactic validity, and similarity-based acceptance; operational cost and execution efficiency through execution time, LLM calls, and token usage; and structural transformation through static code metrics.
Results. Coordination Strategy is the dominant architectural factor: Static configurations achieve substantially higher output validity and lower operational cost across all token-based and call-based measures. Communication Structure exhibits a heterogeneous effect, strongly influencing call-based behavior but having limited impact on token consumption. The two factors interact significantly on all cost measures, with the strongest interaction observed on LLM calls. Among the evaluated configurations, Hierarchical+Static consistently provides the best tradeoff between validity and cost, whereas Centralized+Dynamic fails to produce acceptable output for a substantial fraction of instances.
Conclusions. Coordination Strategy exerts more consistent effects than Communication Structure on both output validity and operational cost, while the interaction between the two factors is central to explaining architectural behavior. The study demonstrates how controlled factorial designs can isolate architectural effects in LLM-MAS research, and it provides empirical guidance for designing cost-efficient architectures with strong output validity for SE tasks.
Cite as
Alexander Puma Pucho and Bruno B. P. Cafeo. A Factorial Comparison of LLM-Based Multi-Agent Architectures for Automated Python Refactoring. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 37:1-37:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{pumapucho_et_al:LIPIcs.ESEM.2026.37,
author = {Puma Pucho, Alexander and Cafeo, Bruno B. P.},
title = {{A Factorial Comparison of LLM-Based Multi-Agent Architectures for Automated Python Refactoring}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {37:1--37:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.37},
URN = {urn:nbn:de:0030-drops-280059},
doi = {10.4230/LIPIcs.ESEM.2026.37},
annote = {Keywords: Architectural Comparison, Code Refactoring, LLM-based Multi-Agent Systems, Software Maintenance Automation, Multi-Agent Architectures}
}
Document
Technical Track Paper
Authors:
Brage Isak Keiserås, Bo Markussen, Ingrid Chieh Yu, Ken Friis Larsen, Maja H. Kirkeby, and Michael Kirkedal Thomsen
Abstract
Background. Compilers are actively developed programs, with newer versions implicitly promising better performance.
Aim. In this paper, we evaluate whether the cost of the expected improvements in optimisations in newer versions of the GCC C compiler pays off, both in terms of execution time and energy consumption.
Method. To do so, we select 10 major versions of GCC (spanning GCC-4 to GCC-13), and four different optimization flags (-Og, -O2, -O3, and -Os) to compile and run benchmarks written in C from the SPEC CPU 2017 benchmarking suite, measuring both the energy and time of compilation and running of the benchmarks. A linear mixed-effects model mapping the GCC major version number to a numerical variable is then fitted to the data.
Result. We find a statistically significant trend of decreasing running costs when optimized with -O2 and -O3, with an overall decrease in costs (for both energy and time) amounting to 13% for -O2 and 9% for -O3 overall (from GCC-4 to GCC-13). However, this is not without a cost, as there is also a statistically significant trend of increasing compilation cost for the optimization levels tested except -Og, with an increase amounting to 17% in energy and 15% in time with -O3 enabled, and 18% in energy and 16% in time with -O2 enabled. Some benchmarks benefit more than others, which we hypothesize is due to more advanced vectorization in newer compiler versions.
Conclusion. GCC is producing lower-cost programs, at the expense of higher compilation cost. Still, the trade-off is well worth it as a single additional execution of the compiled program offsets the increased cost of compilation for both the run-time and energy cost.
Cite as
Brage Isak Keiserås, Bo Markussen, Ingrid Chieh Yu, Ken Friis Larsen, Maja H. Kirkeby, and Michael Kirkedal Thomsen. Does GCC Compiler Development Pay for Itself? GCC Versions' Impact on Energy Consumption. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 38:1-38:19, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{keiseras_et_al:LIPIcs.ESEM.2026.38,
author = {Keiser\r{a}s, Brage Isak and Markussen, Bo and Yu, Ingrid Chieh and Larsen, Ken Friis and Kirkeby, Maja H. and Thomsen, Michael Kirkedal},
title = {{Does GCC Compiler Development Pay for Itself? GCC Versions' Impact on Energy Consumption}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {38:1--38:19},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.38},
URN = {urn:nbn:de:0030-drops-280064},
doi = {10.4230/LIPIcs.ESEM.2026.38},
annote = {Keywords: Compilers, Compiler Optimisation, GCC, Energy Consumption, Energy Measurement}
}
Document
Technical Track Paper
Authors:
Eduardo Barros, Márcio Ribeiro, Valeria Pontillo, Luana Martins, Fabio Palomba, Baldoino Fonseca, Rohit Gheyi, and Ivan Machado
Abstract
Background. Infrastructure-as-Code (IaC) has become a standard in software deployment, with Dockerfiles among its most prevalent formats. However, Dockerfiles often exhibit Dockerfile smells, violations of best practices that can harm reproducibility, build time, image size, and security. Although prior work has proposed detectors, fixes, and refactoring catalogs, existing knowledge remains mostly descriptive or illustrative, lacking explicit transformation rules with source and target structures, and applicability conditions intended to support behavior-preserving application.
Aims. This work aims to provide a catalog of formalized transformation rules for removing Dockerfile smells and to evaluate how practitioners perceive the resulting Dockerfile.
Method. We formalized rules for a set of ten Dockerfile transformations and conducted a mixed-method empirical study with 59 participants: 39 survey respondents and 20 interviewees from industry and academia with varied levels of Docker experience.
Results. Participants generally preferred the resulting Dockerfiles, which achieved an overall mean acceptance rate of 76%. Nine rules were accepted by a majority of participants, with the highest acceptance rate reaching 95%. Participants favored transformations that improved stability, security, and build time, but were less favorable toward transformations that reduced readability or made debugging harder.
Conclusions. Our findings suggest that formalized Dockerfile transformations are a promising basis for refactoring support, but their adoption depends on balancing technical quality improvements with developers’ practical expectations.
Cite as
Eduardo Barros, Márcio Ribeiro, Valeria Pontillo, Luana Martins, Fabio Palomba, Baldoino Fonseca, Rohit Gheyi, and Ivan Machado. A Catalog of Transformations to Remove Dockerfile Smells. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 39:1-39:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{barros_et_al:LIPIcs.ESEM.2026.39,
author = {Barros, Eduardo and Ribeiro, M\'{a}rcio and Pontillo, Valeria and Martins, Luana and Palomba, Fabio and Fonseca, Baldoino and Gheyi, Rohit and Machado, Ivan},
title = {{A Catalog of Transformations to Remove Dockerfile Smells}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {39:1--39:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.39},
URN = {urn:nbn:de:0030-drops-280070},
doi = {10.4230/LIPIcs.ESEM.2026.39},
annote = {Keywords: Docker, Dockerfile, Refactoring, Catalog, Transformation, Quality}
}
Document
Technical Track Paper
Authors:
Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, and Fabio Santos
Abstract
Background. Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. While Large Language Model (LLM)-based APR systems have shown promise, prior studies primarily focus on overall repair effectiveness. The effects of bug complexity, fault localization, reasoning settings, and repair cost-effectiveness remain insufficiently explored.
Aims. This study presents a comprehensive empirical analysis of LLM-based APR, focusing on how repair performance is shaped by bug complexity, fault localization, reasoning settings, and costs.
Method. We construct a curated dataset by collecting algorithmic bugs from AtCoder, a competitive programming platform. We evaluate two APR techniques (ChatRepair and CodeCorrector) using three LLMs (DeepSeek, GPT, and Llama), with multiple model variants and reasoning settings, and examine their performance across diverse levels of bug complexity and localization strategies through a multi-dimensional empirical framework and statistical analysis.
Results. Although structurally complex bugs and imprecise fault localization make repair more challenging, LLM-based APR techniques still achieve competitive repair effectiveness. Imprecise fault localization can substantially enlarge the performance gap between APR techniques. Furthermore, higher-cost LLMs and stronger reasoning settings do not consistently yield better cost-efficiency, revealing a nontrivial trade-off between repair effectiveness and computational cost. We further observe that the impact of reasoning strategies varies considerably across different LLM families, affecting both repair effectiveness and cost-efficiency.
Conclusions. Over 50% of moderately complex bugs can be repaired by low-cost LLM-based APR techniques. The repair effectiveness gap between APR techniques becomes larger as fault localization becomes less precise. GPT-5 repairs 7 and 39 more complex bugs than DeepSeek-V4-pro and DeepSeek-V3.2, respectively; whereas the total repair cost of DeepSeek-V3.2 across the non-reasoning and reasoning stages is only approximately 4.6% and 7.7% of GPT-5 and DeepSeek-V4-pro, respectively.
Cite as
Junchi Liu, Ali Bigdeli, Roya Daneshi, Atu Ambala, Sudipto Ghosh, and Fabio Santos. Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 40:1-40:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{liu_et_al:LIPIcs.ESEM.2026.40,
author = {Liu, Junchi and Bigdeli, Ali and Daneshi, Roya and Ambala, Atu and Ghosh, Sudipto and Santos, Fabio},
title = {{Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-Efficiency}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {40:1--40:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.40},
URN = {urn:nbn:de:0030-drops-280081},
doi = {10.4230/LIPIcs.ESEM.2026.40},
annote = {Keywords: Automated Program Repair, Large Language Models, Fault Localization, Bug Complexity, Cost-efficiency}
}
Document
Technical Track Paper
Authors:
Hui Sun, Anderson Uchôa, Rohit Gheyi, and Wesley K. G. Assunção
Abstract
Background. Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to believe that they can reason about program semantics. A fundamental test of this capability is assessing code functional equivalence. However, existing evaluations primarily focus on single-language settings or rely on synthetically generated code, raising concerns about whether current results reflect true semantic understanding.
Aims. We investigate whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity.
Method. We introduce PolyHuman, a dataset of human-written programs in CPP, Java, and Python. Using this dataset, we evaluate intra- and inter-language equivalence detection across open-weight and proprietary LLMs, selecting GPT-o4-mini as a representative model to assess stability. We then manually analyze 81 cases of systematic disagreement in which models incorrectly judge functional equivalence, examining the code logic and the generated Chain-of-Thought reasoning. Finally, we categorize these failures and compare them across GPT-o4-mini, Claude-Opus-4.7, and Gemini-3-Flash to determine whether they reflect model-specific issues or broader limitations of state-of-the-art LLMs.
Results. We identify a difficulty-dependent breakdown in equivalence judgment (harder problems make the model increasingly prone to misclassifying non-equivalent code as equivalent), a model-specific sensitivity to programming language for the best-performing model (particularly a more conservative behavior on Python), and a partial reliance on similarity-based cues. GPT-o4-mini also shows substantial run-to-run instability under identical settings, indicating inconsistent rather than absent capability. Its failures are mainly due to (1) a lack of knowledge relevant to judging code functional equivalence and (2) abstraction-level reasoning failure, such as premature conclusion, semantic-level misunderstanding, and implementation-level error. We also observed that several analyzed failures involved loops and conditional branches.
Conclusions. Current LLMs do not reliably capture functional equivalence within or across languages. Our findings highlight the need for more realistic benchmarks and suggest improvements through better reasoning-path selection and by addressing reasoning gaps across the identified failure levels.
Cite as
Hui Sun, Anderson Uchôa, Rohit Gheyi, and Wesley K. G. Assunção. Evaluating Language Models on Cross-Language Code Functional Equivalence. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 41:1-41:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{sun_et_al:LIPIcs.ESEM.2026.41,
author = {Sun, Hui and Uch\^{o}a, Anderson and Gheyi, Rohit and Assun\c{c}\~{a}o, Wesley K. G.},
title = {{Evaluating Language Models on Cross-Language Code Functional Equivalence}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {41:1--41:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.41},
URN = {urn:nbn:de:0030-drops-280091},
doi = {10.4230/LIPIcs.ESEM.2026.41},
annote = {Keywords: Empirical software engineering, code functional equivalence, generative AI, cross-language analysis, benchmark, failure analysis}
}
Document
Technical Track Paper
Authors:
Negar Alizadeh, Nishant Saurabh, and Fernando Castor
Abstract
Background. Code completion is one of the most widely used applications of large language models (LLMs) in software development. In addition to proprietary coding assistants, powerful open-weight LLMs are increasingly adopted for locally deployed code completion systems, partly motivated by privacy concerns. Despite advances in LLM accuracy, the energy cost of inference in code completion tasks remains underexplored, particularly under large-context workloads and across programming languages.
Aims. This study investigates the trade-off between accuracy and energy consumption in LLM-based code completion and analyzes how workload characteristics, context size, and model scale influence inference energy usage.
Method. We evaluate 25 open-weight LLMs on two complementary code completion workloads: repository-level next-line completion (left-to-right prediction) with varying context sizes on the RepoBench dataset, and fill-in-the-middle (FIM) code completion across Python, Java, and Rust on the McEval dataset. We further analyze the influence of input tokens, output tokens, model size, and their interactions on energy consumption using correlation analysis and cluster-robust linear regression models.
Results. Our findings show that the dominant drivers of energy consumption depend strongly on the structure of the completion task. In RepoBench, energy consumption is primarily influenced by input context size and its interaction with model scale, whereas in McEval, output generation and its interaction with active parameter count become the dominant factors. We further observe that output generation is substantially more energy-intensive per token than prompt processing. Across both benchmarks, smaller and heavily quantized models frequently achieve Pareto-optimal trade-offs, often providing accuracy comparable to larger FP16 models while consuming substantially less energy.
Conclusions. The energy behavior of LLM inference in software engineering tasks depends strongly on workload structure, context length, and model scale. Our results suggest that increasing model size or context length does not necessarily lead to proportionally better completion quality, while quantization can substantially improve energy efficiency with limited accuracy degradation. These findings contribute toward more energy-aware deployment strategies for sustainable AI-assisted software development.
Cite as
Negar Alizadeh, Nishant Saurabh, and Fernando Castor. Green AI: Cost of LLM-Based Code Completion. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 42:1-42:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{alizadeh_et_al:LIPIcs.ESEM.2026.42,
author = {Alizadeh, Negar and Saurabh, Nishant and Castor, Fernando},
title = {{Green AI: Cost of LLM-Based Code Completion}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {42:1--42:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.42},
URN = {urn:nbn:de:0030-drops-280104},
doi = {10.4230/LIPIcs.ESEM.2026.42},
annote = {Keywords: Green AI, Green LLMs, Code Completion, Coding Assistant, Energy Efficiency, Trade-Offs, Software Development, Model Quantization}
}
Document
Technical Track Paper
Authors:
Xing Cui, Jingzheng Wu, Tianyue Luo, and Xiang Ling
Abstract
Background. The sustainable evolution of Open Source Software (OSS) depends not only on code production but also on continuous knowledge transfer and collaborative support within communities. However, existing contribution metrics primarily center on code commits and merges, overlooking the value of non-coding activities embedded in technical discussions. Among these activities, Non-explicit Mentoring (NEM) refers to guidance-oriented interactions in Pull Request (PR) and Issue discussions, where contributors convey technical knowledge, facilitate problem diagnosis, clarify design rationale, provide normative guidance, and share feedback within natural conversational contexts. Although prior studies acknowledge the existence of such behaviors, systematic quantitative investigations of their scale, structural patterns, and impact remain limited due to reliance on qualitative methods, insufficient automated detection techniques, and the scarcity of high-quality annotated datasets.
Aims. To address these challenges, this paper proposes MentoScope, a large language model (LLM)-based automated framework for fine-grained identification and classification of NEM in GitHub collaboration contexts.
Method. Built upon Llama-3.1-8B, MentoScope undergoes a three-stage training pipeline. First, continual pre-training (CPT) with LoRA is applied on 92,778 PRs and 118,569 Issues for OSS domain adaptation. Second, supervised fine-tuning (SFT) on 15,214 high-quality annotated samples enables recognition of 8 predefined NEM categories. Third, Odds Ratio Preference Optimization (ORPO) with 3,280 preference pairs aligns model outputs with human judgment to enhance classification robustness.
Results. MentoScope achieves F1-scores of 94.20% and 80.13% on NEM identification and classification tasks respectively, significantly outperforming all baselines, with ablation studies confirming the necessity of each training stage. Large-scale analysis of 591,154 comments from 500 OSS projects reveals that NEM is present in over 70% of collaborative interactions, with distributional patterns varying systematically across project scales and discussion contexts. Furthermore, survival analysis on 1,245 newcomer contributors shows that NEM recipients achieve a 12.6% improvement in PR merge rates and a 35% extension in average active tenure compared to the control group.
Conclusion. These findings demonstrate that NEM constitutes a prevalent and impactful collaborative behavior in open source communities, with significant implications for newcomer integration and community sustainability.
Cite as
Xing Cui, Jingzheng Wu, Tianyue Luo, and Xiang Ling. The Overlooked Spirit of Open Source: LLM-Based Analysis of Non-Explicit Mentoring in OSS Communities. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 43:1-43:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{cui_et_al:LIPIcs.ESEM.2026.43,
author = {Cui, Xing and Wu, Jingzheng and Luo, Tianyue and Ling, Xiang},
title = {{The Overlooked Spirit of Open Source: LLM-Based Analysis of Non-Explicit Mentoring in OSS Communities}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {43:1--43:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.43},
URN = {urn:nbn:de:0030-drops-280119},
doi = {10.4230/LIPIcs.ESEM.2026.43},
annote = {Keywords: open source software, non-explicit mentoring, large language models, contributor retention}
}
Document
Technical Track Paper
Authors:
Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, and Li Zhang
Abstract
Background. LLM coding assistants are widely used in development, but evidence about their reliability mostly comes from mainstream languages such as Python and Java. Their behaviour on emerging languages, where public code and documentation are scarce, is far less clear.
Aims. We ask how capable current LLMs are on Cangjie, an emerging language in the HarmonyOS ecosystem, and which assistive technique is worth its token cost.
Method. We build CangjieBench, 248 manually translated tasks derived from HumanEval and ClassEval, run under Text-to-Code and Code-to-Code modalities. We evaluate six LLMs under five methods: direct generation, syntax-constrained prompting, retrieval over documentation, retrieval over code, and agent-based workflows. We report Pass@1, compile rate, token cost, and a seven-category failure taxonomy validated by two LLM annotators (κ ≥ 0.98).
Results. Direct generation is unusable on most models, and over four-fifths of failed samples carry foreign syntax that the Cangjie compiler rejects. A 2,146-token syntax reference gives the best accuracy-to-cost trade-off among prompt-based methods. The best agent configuration reaches the highest Pass@1 at 10-120x the token cost of a single prompt, while weaker agents cannot match a syntax-constrained call on the same backbone. A Python reference does not always help: on Code-to-Code it can pull models toward Python idioms and drop compile rates below Text-to-Code.
Conclusions. No single method dominates both accuracy and token cost. A practical assistant for an emerging language should ship a concise syntax cheat sheet first, escalate to curated retrieval for API-heavy tasks, and reserve agent loops for tasks that need multi-turn repair. The replication package is at https://github.com/cjhCoder7/CangjieBench.
Cite as
Junhang Cheng, Fang Liu, Jia Li, Chengru Wu, Nanxiang Jiang, and Li Zhang. Can LLM Coding Assistants Support Emerging Programming Languages? An Empirical Study on Cangjie. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 44:1-44:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{cheng_et_al:LIPIcs.ESEM.2026.44,
author = {Cheng, Junhang and Liu, Fang and Li, Jia and Wu, Chengru and Jiang, Nanxiang and Zhang, Li},
title = {{Can LLM Coding Assistants Support Emerging Programming Languages? An Empirical Study on Cangjie}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {44:1--44:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.44},
URN = {urn:nbn:de:0030-drops-280128},
doi = {10.4230/LIPIcs.ESEM.2026.44},
annote = {Keywords: Large language models, Code generation, Code translation, Emerging programming languages, Empirical software engineering, Benchmark, Cangjie}
}
Document
Technical Track Paper
Authors:
Shawal Khalid, Joaquin Tuckett, and Chris Brown
Abstract
Background. Smart contract vulnerabilities pose significant security risks in blockchain systems. Automated static auditing tools are commonly used in practice, yet their effectiveness varies across vulnerability categories and at scale. Recently, large language model (LLM)-based tools have been proposed, but their performance is rarely evaluated against real-world contracts and alongside established analyzers.
Aim. This paper presents a systematic benchmarking study of smart contract auditing tools under a single, explicitly defined task: identifying smart contract vulnerabilities in Solidity contracts.
Method. We evaluate static and LLM-based tools (n = 7) on two benchmarks: SmartBugs Curated (n = 143) and a filtered FORGE subset of real-world audit-derived contracts (n = 173). We use a tool-agnostic evaluation pipeline to normalize heterogeneous tool outputs, reporting accuracy, top-K detection, execution robustness, and cross-dataset performance.
Results. Our findings expose systematic trade-offs across auditing approaches. Static tools optimized for broad vulnerability coverage tend to over-report categories, potentially increasing triage burden, while exhibiting stability limitations on real-world contracts. LLM-driven auditors demonstrate strong coverage across many vulnerability categories, but similarly suffer from over-prediction.
Conclusions. These results show that current smart contract auditing tools differ less in their ability to identify potential vulnerabilities than in their ability to do so precisely and reliably, motivating future auditing workflows for practical smart contract vulnerability detection.
Cite as
Shawal Khalid, Joaquin Tuckett, and Chris Brown. Do Smart Contract Auditing Results Transfer Across Datasets? A Two-Benchmark Empirical Study of Static and LLM-Based Security Tools. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 45:1-45:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{khalid_et_al:LIPIcs.ESEM.2026.45,
author = {Khalid, Shawal and Tuckett, Joaquin and Brown, Chris},
title = {{Do Smart Contract Auditing Results Transfer Across Datasets? A Two-Benchmark Empirical Study of Static and LLM-Based Security Tools}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {45:1--45:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.45},
URN = {urn:nbn:de:0030-drops-280137},
doi = {10.4230/LIPIcs.ESEM.2026.45},
annote = {Keywords: Smart contract vulnerability detection, static analysis, symbolic execution, benchmarking, large language models}
}
Document
Technical Track Paper
Authors:
Vinicius Carvalho Lopes, Matthew Zampino, Christian Garcia, Joanna C. S. Santos, and Adriana Sejfia
Abstract
Background. Vulnerability-Contributing Commits (VCCs) are the code changes at which a software system turns from safe to unsafe with respect to a specific vulnerability. Accurate VCC identification is crucial for training vulnerability detection models, yet the reliability of existing VCC datasets remains underexplored. If these VCC datasets contain mislabeled commits, models trained on them may learn incorrect patterns.
Aims. We present a two-part study of VCC dataset quality: characterizing how VCCs are defined and identified across the literature, and empirically measuring the accuracy of existing VCC datasets.
Method. We review 38 papers and, based on their findings, propose an operational VCC definition based on security state transitions and testable through parent commit validation. We apply this definition in an exploit-based validation study on 6 datasets aggregating 3,105 VCC samples, of which 168 (5.41%) have publicly documented exploits enabling empirical validation.
Results. We found significant terminological and definitional inconsistencies: 8 distinct terms with varying definitions, 15.8% of papers lacking explicit definitions, and confusion between commits that contain versus commits that introduce vulnerabilities. Among the 168 validated samples, we observe a 79.2% false positive rate, rising to 88.4% in the subset whose verdicts rest entirely on runtime exploit reproduction, with 47.0% being commits labeled as VCCs although the vulnerability was already exploitable in the parent commit.
Conclusions. Existing VCC datasets contain labeling errors detectable through exploit-based parent commit validation. Our findings call for standardized VCC definitions and parent commit validation to distinguish VCC commits from those that merely modify already-vulnerable code.
Cite as
Vinicius Carvalho Lopes, Matthew Zampino, Christian Garcia, Joanna C. S. Santos, and Adriana Sejfia. On the Quality of Vulnerability-Contributing Commit Datasets: A Systematic Review and Exploit-Based Evaluation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 46:1-46:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{carvalholopes_et_al:LIPIcs.ESEM.2026.46,
author = {Carvalho Lopes, Vinicius and Zampino, Matthew and Garcia, Christian and Santos, Joanna C. S. and Sejfia, Adriana},
title = {{On the Quality of Vulnerability-Contributing Commit Datasets: A Systematic Review and Exploit-Based Evaluation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {46:1--46:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.46},
URN = {urn:nbn:de:0030-drops-280146},
doi = {10.4230/LIPIcs.ESEM.2026.46},
annote = {Keywords: vulnerability-contributing commits, software security, dataset quality, systematic literature review, exploit validation}
}
Document
Technical Track Paper
Authors:
Jiayu Liu and Yanjun Wu
Abstract
Background. Targeted unit tests, which aim to reveal a bug at a specific location in the program, are useful for various software development tasks, such as static analysis alarm validation, crash reproduction, or patch testing. While automated unit test generation has progressed significantly, existing approaches remain inefficient for targeted tests due to their reliance on manually-crafted heuristics. To address this shortcoming, Large Language Model (LLM) agents emerge as a potent solution that is free from predefined heuristics, owing to their superior reasoning and tool-use capabilities. Moreover, prior studies have shown that LLM agents can achieve comparable performance across several software engineering tasks with only minor adaptations. However, whether this adaptability extends to targeted unit test generation remains unexplored.
Aims. This study aims to investigate the capabilities and limitations of LLM agents in targeted unit test generation under minimal custom setup.
Method. We evaluate various combinations of frontier agent scaffolds and models on a benchmark featuring 210 real-world Java runtime exception bugs across 66 software projects. Specifically, LLM agents are tasked with generating a unit test suite that reproduce a target bug given only its location and the corresponding codebase.
Results. Our evaluation shows that the top-performing combinations achieve a 89.5% success rate in target bug reproduction, surpassing the state-of-the-art unit test generation tools. Nevertheless, we identify recurring failure modes that illustrate current LLM agents' blind spots.
Conclusions. These findings underscore the potential of LLM agents in targeted unit test generation, and provide valuable insights for further research in this area.
Cite as
Jiayu Liu and Yanjun Wu. Beyond Heuristics? Rethinking Targeted Unit Test Generation in the Era of LLM Agents. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 47:1-47:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{liu_et_al:LIPIcs.ESEM.2026.47,
author = {Liu, Jiayu and Wu, Yanjun},
title = {{Beyond Heuristics? Rethinking Targeted Unit Test Generation in the Era of LLM Agents}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {47:1--47:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.47},
URN = {urn:nbn:de:0030-drops-280156},
doi = {10.4230/LIPIcs.ESEM.2026.47},
annote = {Keywords: Large Language Model, Agents, Unit Test Generation, Empirical Study}
}
Document
Technical Track Paper
Authors:
Hamza Attarwala, Moataz Chouchen, Mohammad Hamdaqa, and Omar Alam
Abstract
Background. Large Language Models (LLMs) are increasingly used to generate Object Constraint Language (OCL) constraints from natural language specifications and UML class diagrams. However, existing work mainly focuses on improving accuracy, with limited understanding of why these models fail.
Aims. This study investigates factors associated with LLM failures in OCL generation by examining the task from a graph-reasoning perspective.
Method. We conduct an empirical evaluation using the PathOCL dataset across six LLMs. We analyze the relationship between OCL correctness and UML structural properties (e.g., navigation depth and model complexity), lexical similarity, prompt ordering strategies, and graph-aware prompting.
Results. We find that, within the evaluated dataset and models, OCL generation correctness is negatively associated with navigation depth and structural complexity. Lexical similarity provides limited explanatory power, while textual ordering is associated with differences in performance. Graph-based prompting yields partial improvements but does not eliminate errors involving structural navigation.
Conclusions. The findings are consistent with structural demands being an important contributor to OCL generation difficulty, while not establishing graph reasoning as the sole or primary cause of failure.
Cite as
Hamza Attarwala, Moataz Chouchen, Mohammad Hamdaqa, and Omar Alam. Why Do LLMs Fail at OCL Generation? A Graph Reasoning Perspective. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 48:1-48:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{attarwala_et_al:LIPIcs.ESEM.2026.48,
author = {Attarwala, Hamza and Chouchen, Moataz and Hamdaqa, Mohammad and Alam, Omar},
title = {{Why Do LLMs Fail at OCL Generation? A Graph Reasoning Perspective}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {48:1--48:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.48},
URN = {urn:nbn:de:0030-drops-280160},
doi = {10.4230/LIPIcs.ESEM.2026.48},
annote = {Keywords: Object Constraint Language (OCL), Unified Modelling Language (UML), Model-Driven Engineering (MDE), Large Language Models (LLM), Graph Reasoning}
}
Document
Technical Track Paper
Authors:
Nils Kiele, Zainab Saad, Zirui Wang, Steve Drew, and Samira Ebrahimi Kahou
Abstract
Background. Mutation testing evaluates test-suite adequacy by injecting synthetic faults into program code. However, traditional rule-based tools often generate large numbers of trivial, redundant, or equivalent mutants that limit their practical use for identifying gaps in a test suite. While recent large language model (LLM)-based approaches generate more realistic faults, most remain test-blind: The model sees only the source code and cannot reason about what existing tests already cover. Ignoring such tests means neglecting context that could help generate higher-quality mutants and thus stronger tests to patch the remaining test suite gaps.
Aims. We propose test-aware mutant generation, in which an LLM receives the problem statement, canonical solution and base tests in a single prompt, and must generate a nontrivial mutant that passes the base unit tests.
Method. We evaluate this approach across a set of five LLMs - Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.1 Codex Mini, GPT 4.1 Mini, Qwen3-32B - on the HumanEval and MBPP benchmarks. The extended EvalPlus test suites serve as an automated oracle to verify whether surviving mutants represent genuine bugs.
Results. Test-aware prompting yields verified fault rates of 87.7% (HumanEval) and 79.1% (MBPP), meaning these mutants pass all base tests but are caught by the oracle. This vastly outperforms the matched test-blind prompting (which yields only 12.2% and 23.0%, respectively) and the traditional rule-based tool mutmut (4.4% and 5.7%). While fault subtlety (the fraction of extended tests a mutant fails) remains comparable across all three methods, test-awareness minimizes the computational cost per verified fault, compared to test-blind prompting.
Conclusions. Exposing LLMs to existing unit tests shifts mutant generation from untargeted bug injection toward effective discovery of weaknesses in an existing test suite. Our work establishes a concrete foundation for future research to scale test-aware mutant generation to production-level environments.
Cite as
Nils Kiele, Zainab Saad, Zirui Wang, Steve Drew, and Samira Ebrahimi Kahou. Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 49:1-49:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{kiele_et_al:LIPIcs.ESEM.2026.49,
author = {Kiele, Nils and Saad, Zainab and Wang, Zirui and Drew, Steve and Ebrahimi Kahou, Samira},
title = {{Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {49:1--49:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.49},
URN = {urn:nbn:de:0030-drops-280176},
doi = {10.4230/LIPIcs.ESEM.2026.49},
annote = {Keywords: Mutation testing, large language models}
}
Document
Technical Track Paper
Authors:
Benjamin Agyekum and Fabio Santos
Abstract
Background. Iterative feedback loops have become the dominant paradigm for improving LLM-generated Infrastructure-as-Code (IaC): validators such as Checkov and terraform validate feed error signals back to the model for successive repair attempts. Prior work reports cumulative-best metrics, which are monotonically non-decreasing by construction, so the raw per-iteration security trajectory has never been examined in the IaC domain.
Aims. We study security regression (a previously-passing CIS Benchmark check that fails after a repair iteration) to determine whether, and how often, iterative LLM repair degrades security while fixing other issues.
Method. We analyze 5,968 scenario timelines from the IaC-Eval benchmark, each one scenario run through one configuration for up to 5 repair iterations. The 15 configurations comprise six model-specific RAG and nine model-aggregated non-RAG configurations, three temperatures each, and together they yield 4,440 iteration transitions with Checkov data on both sides. We track 30 individual CIS check IDs and classify regression root causes from code diffs, under two detection modes: standard (inclusive) and strict (exclusive check failures only).
Results. Under standard (inclusive) detection, 13.8% of scenarios (24.8% of transitions) exhibit at least one regression. Under strict detection, which counts only unambiguous, exclusive check failures, the rate falls to 3.3% of scenarios (5.2% of transitions). This gap indicates that most apparent regressions are multi-resource measurement artifacts rather than genuine exclusive failures. Resource restructuring (79.0%) is the dominant root cause. Regression transitions show 2.6× more code churn (Cohen’s d = 0.90) and 4.9× higher strict-mode check volatility (d = 1.49). Of standard-mode regressions, 36.6% self-correct within an average of 1.2 iterations, and iteration 3 is the optimal stopping point.
Conclusions. Iterative IaC repair does introduce security regressions, but most apparent regressions are multi-resource measurement artifacts. The conservative, defensible rate is approximately 3.3% of scenarios. Our findings motivate security-aware feedback-loop design and provide actionable iteration-budget guidance.
Cite as
Benjamin Agyekum and Fabio Santos. Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-As-Code Repair. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 50:1-50:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{agyekum_et_al:LIPIcs.ESEM.2026.50,
author = {Agyekum, Benjamin and Santos, Fabio},
title = {{Does Fixing Break Security? An Empirical Study of Security Degradation in Iterative LLM-Driven Infrastructure-As-Code Repair}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {50:1--50:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.50},
URN = {urn:nbn:de:0030-drops-280186},
doi = {10.4230/LIPIcs.ESEM.2026.50},
annote = {Keywords: Infrastructure as Code, Security Regression, LLM Code Repair, Terraform, CIS Compliance, Iterative Feedback}
}
Document
Technical Track Paper
Authors:
Gyumin Nam and Geunseok Yang
Abstract
Background. Regression testing has traditionally focused on source-code changes and well-specified software dependencies such as compiler versions, library versions, and operating-system platforms. However, LLM-enabled code-generation systems also depend on non-code artifacts (the model, the prompt template, and the retrieval corpus) that can change independently of the application code. Updates to any of these artifacts can silently regress previously passing tasks, and the lightweight partial-test oracles commonly used in continuous evaluation may not detect such failures. Despite the practical relevance, paired regression rates, measurement validity, and escaped-failure rates under controlled non-code dependency changes have received little empirical attention.
Aims. This work empirically analyzes how changes to models, prompts, and retrieval data affect the behavior of LLM-enabled code-generation systems, framing the problem as dependency-aware regression testing rather than a model or retrieval comparison. The questions addressed are failure-rate sensitivity, paired-regression frequency under the answer-bearing canonical retrieval condition, the measurement-validity decomposition of canonical regression signals (answer-bearing baseline rescue vs. non-canonical retrieval sensitivity), and how often generations pass a lightweight oracle yet fail a stronger oracle, reported over all generations and, conditionally, among stronger-oracle failures.
Method. We designed a controlled empirical study that treats models, prompts, and retrieval data as software dependencies of an LLM-enabled code-generation pipeline. Starting from a fixed baseline, we varied one dependency dimension at a time across three open-weight 7B-class code LLMs, three prompt templates, and four retrieval-data conditions (one being a synthetic stress test), evaluated on 120 CodeRAG-Bench HumanEval tasks under three random seeds (2,880 generations across eight conditions). For each (task, condition, seed) triple, we recorded syntactic validity and per-assertion execution outcomes to derive failure rate, paired regression rate, escaped-failure rate, and across-seed flakiness, and compared each changed condition to the baseline using paired McNemar tests.
Results. In the answer-bearing canonical benchmark condition, removing the retrieved document regressed 11.67% of previously passing tasks (D₁, p = 1.22 × 10⁻⁴, stable under task-level Bonferroni correction). Replacing it with a different-task mismatched document regressed 9.17% (D₂, p = 9.77 × 10⁻⁴, significant before correction but not after; reported as an auxiliary stressor). A copy-template audit showed that the M₀ baseline output was a normalized byte-level copy of the canonical retrieved document in 331 of 360 generations, and 12 of 14 D₁-regressed tasks exhibited this property. A title-based self-excluded BM25 proxy and an answer-stripped D₀ check showed that most of this canonical signal was due to answer-bearing baseline rescue (BM25 D₁: 3 regressions / 3 recoveries; D₂: 3 / 4); thus, harmful non-canonical retrieval degradation was not established in this sample. The intentionally permissive O₁ probe let 36/360 D₁ and 26/360 D₂ generations pass despite failing O₂; an order-dependent first-three-substantive variant reduced these counts to 8/360 and 3/360. Model and prompt perturbations are reported as framework coverage checks rather than causal claims.
Conclusions. The strongest apparent regression signal arose from the answer-bearing canonical retrieval condition, but validity analyses showed that this signal primarily reflected a benchmark-induced copy-template dependency rather than an established deployment retrieval risk. These findings support dependency-aware paired regression testing and layered oracles, while cautioning that canonical answer-bearing retrieval regression rates should be interpreted as benchmark-specific stress signals rather than deployment-risk estimates.
Cite as
Gyumin Nam and Geunseok Yang. Hidden Dependencies in LLM-Enabled Code Generation: An Empirical Study of Regressions and Escaped Failures. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 51:1-51:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{nam_et_al:LIPIcs.ESEM.2026.51,
author = {Nam, Gyumin and Yang, Geunseok},
title = {{Hidden Dependencies in LLM-Enabled Code Generation: An Empirical Study of Regressions and Escaped Failures}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {51:1--51:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.51},
URN = {urn:nbn:de:0030-drops-280195},
doi = {10.4230/LIPIcs.ESEM.2026.51},
annote = {Keywords: LLM-enabled code generation, regression testing, retrieval-augmented generation, test oracles, empirical software engineering}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Luisa Bartl, Marvin Wyrich, Sven Apel, and Janet Siegmund
Abstract
Code comprehension is an inevitable part of the software development process and a frequent target of empirical research. However, the field lacks a shared definition of the construct. Studies operationalize code comprehension in diverse ways, often without explicitly stating what aspect of understanding they measure. As a result, researchers design study-specific tasks and measures, limiting comparability across studies and hindering cumulative knowledge building.
To initiate a coordinated approach to assess code comprehension, we introduce a facet-based perspective that distinguishes between analytical skills, abstraction, and critical evaluation. We demonstrate the benefit of explicitly stating the targeted facet(s) of code comprehension and aligning task design accordingly. Building on this perspective, we propose a method to guide researchers from research goals to concrete tasks and study design, including considerations for question formulation and response formats. Rather than immediately prescribing a single standardized instrument, our approach takes a first step toward this long-term goal by enabling researchers to design adaptable yet conceptually grounded measurements. By making operationalizations explicit and tasks reusable, we initiate a cumulative process through which the community can iteratively converge on standardized instruments and, ultimately, a shared understanding of code comprehension.
Cite as
Luisa Bartl, Marvin Wyrich, Sven Apel, and Janet Siegmund. A Facet-Based Perspective on Code Comprehension Task Design for Empirical Research. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 52:1-52:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{bartl_et_al:LIPIcs.ESEM.2026.52,
author = {Bartl, Luisa and Wyrich, Marvin and Apel, Sven and Siegmund, Janet},
title = {{A Facet-Based Perspective on Code Comprehension Task Design for Empirical Research}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {52:1--52:15},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.52},
URN = {urn:nbn:de:0030-drops-280208},
doi = {10.4230/LIPIcs.ESEM.2026.52},
annote = {Keywords: Code comprehension, Construct validity, Empirical software engineering}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Lukas Mehl
Abstract
Migrating open-source systems to post-quantum cryptography requires knowing where classical public-key primitives still appear in code, yet much empirical work has targeted dependency graphs or misuse of APIs rather than primitive-level, multi-language usage at scale. This paper describes a static analyzer that scans Python, Java, and Go repositories using syntax-aware parsing, classifies cryptographic calls and imports as post-quantum-vulnerable, quantum-safe, or PQC-ready, and aggregates findings for ecosystem-level measurement.
Applied to a stratified sample of about fourteen and a half thousand GitHub repositories, vulnerable primitives are found to be widespread at the level of explicit calls and imports, with strongly language-dependent rates shaped in part by how transport and certificate stacks are treated in static analysis. Evidence of PQC-ready APIs is exceedingly rare by comparison, underscoring a stark adoption gap relative to migration goals. Methodological limitations are stated so the baseline can be interpreted, reproduced, and extended.
Cite as
Lukas Mehl. PQC-Readiness of Open-Source Software: An Empirical Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 53:1-53:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{mehl:LIPIcs.ESEM.2026.53,
author = {Mehl, Lukas},
title = {{PQC-Readiness of Open-Source Software: An Empirical Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {53:1--53:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.53},
URN = {urn:nbn:de:0030-drops-280211},
doi = {10.4230/LIPIcs.ESEM.2026.53},
annote = {Keywords: post-quantum cryptography, static analysis, empirical study, open-source software, cryptographic primitives}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Muhammad Auwal Abubakar, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Jie M. Zhang, Christoph Treude, Sebastian Baltes, and Matthias Galster
Abstract
Repository-level configuration artifacts allow developers to provide guidance for agentic AI coding tools, e.g., Claude Code, Gemini. Although prior research has examined repository-shared context files that capture project-level instructions and conventions (e.g., AGENTS.md files), little is known about more task-oriented artifacts such as Agent Plans. We present an exploratory study of Agent Plans in open-source software repositories, examining how plan files are preserved, which development activities they support, and what information they provide to guide agent execution. We screened 36,710 GitHub repositories belonging to engineered software projects and identified 85 Markdown plan files from 10 repositories. Within this concentrated corpus, Agent Plans supported several kinds of software engineering work, including maintenance, design, construction, quality-related work, and process support. They also provided task-oriented execution guidance, most commonly through implementation steps, concrete files and locations, and testing and validation information. Overall, repository-preserved Agent Plans in these tool-specific directories appear to be a narrow but informative artifact for studying task intent and execution guidance in human-agent workflows.
Cite as
Muhammad Auwal Abubakar, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Jie M. Zhang, Christoph Treude, Sebastian Baltes, and Matthias Galster. An Exploratory Study of Agent Plans for Agentic AI Coding Tools in Open-Source Software. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 54:1-54:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{abubakar_et_al:LIPIcs.ESEM.2026.54,
author = {Abubakar, Muhammad Auwal and Mohsenimofidi, Seyedmoein and Lulla, Jai Lal and Zhang, Jie M. and Treude, Christoph and Baltes, Sebastian and Galster, Matthias},
title = {{An Exploratory Study of Agent Plans for Agentic AI Coding Tools in Open-Source Software}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {54:1--54:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.54},
URN = {urn:nbn:de:0030-drops-280222},
doi = {10.4230/LIPIcs.ESEM.2026.54},
annote = {Keywords: Agentic AI Coding Tools, Configuration Mechanisms, Plans}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Mohamed Haady Tiemtore and Frédéric Loulergue
Abstract
Infrastructure-as-Code (IaC) artefacts such as Puppet manifests, Ansible playbooks, Kubernetes manifests, Helm charts, Terraform configurations, and CloudFormation templates are increasingly used to manage critical infrastructures. Security defects in these artefacts may therefore be propagated rapidly through automated deployment pipelines. This paper reports emerging results from a focused literature review of tools for detecting security issues in IaC artefacts. The current review protocol covers the ACM Digital Library and IEEE Xplore and includes 23 primary studies. We identify the main families of tools captured by this protocol, classify their detection approaches, and analyse how they are evaluated. The review corpus includes rule-based linters, deeper static analyses, graph- and model-based approaches, comparative scanner studies, deployment-oriented analyses, and recent machine-learning and LLM techniques.
Cite as
Mohamed Haady Tiemtore and Frédéric Loulergue. Tools for Detecting Security Issues in Infrastructure-As-Code: A Focused Literature Review. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 55:1-55:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{tiemtore_et_al:LIPIcs.ESEM.2026.55,
author = {Tiemtore, Mohamed Haady and Loulergue, Fr\'{e}d\'{e}ric},
title = {{Tools for Detecting Security Issues in Infrastructure-As-Code: A Focused Literature Review}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {55:1--55:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.55},
URN = {urn:nbn:de:0030-drops-280238},
doi = {10.4230/LIPIcs.ESEM.2026.55},
annote = {Keywords: Infrastructure-as-Code, security tooling, misconfiguration detection, static analysis, empirical software engineering, focused literature review}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Dmytro Polishchuk
Abstract
Quality-oriented mining software repositories studies require more than access to commits, issues, pull requests and CI logs. They require a reproducible way to represent how technical, social, temporal and defect-related evidence is linked across repository history. Existing MSR infrastructures provide important support for repository mining, but they do not directly expose a project-level, time-aware empirical model in which cross-artifact quality evidence can be queried, inspected and reused across alternative study designs.
This paper presents QModel, an emerging infrastructure for reproducible GitHub mining in empirical software quality studies. QModel stores Git, GitHub, CI, timeline, reaction, graph, churn and SZZ-style candidate-defect evidence in a project-centric relational schema. Its SQL-based compilation layer allows researchers to define feature-target datasets over arbitrary analysis units, including pull requests, issues, commits, files, time windows and custom cross-artifact objects.
We evaluate QModel on ansible/ansible and facebook/react, mining 76,475 commits, 70,509 pull requests, 45,908 issues, 2,174,020 timeline events, 315,235 file-change records and 446,161 CI records. The early results show that QModel can materialize analysis-ready datasets with computable target, graph, churn and provenance features, while also making evidence gaps explicit when historical links, CI records, or fixing evidence are incomplete. These results support QModel as an empirical infrastructure for making quality-oriented repository operationalizations explicit, reusable and systematically evaluable.
Cite as
Dmytro Polishchuk. QModel: A Time-Aware GitHub Mining Framework for Empirical Software Quality Studies. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 56:1-56:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{polishchuk:LIPIcs.ESEM.2026.56,
author = {Polishchuk, Dmytro},
title = {{QModel: A Time-Aware GitHub Mining Framework for Empirical Software Quality Studies}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {56:1--56:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.56},
URN = {urn:nbn:de:0030-drops-280245},
doi = {10.4230/LIPIcs.ESEM.2026.56},
annote = {Keywords: Mining software repositories, CI/CD, software quality, reproducibility}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Nicolás E. Díaz Ferreyra, Manish Mahesh Kumar, Nohemí Villarreal, Pankaj Pantel, Immo Brueggemann, and Riccardo Scandariato
Abstract
Threat modeling remains a central task in secure software engineering, as it enables the identification of security issues from system architectures. As Generative Artificial Intelligence (GenAI) becomes increasingly pervasive across software systems, traditional threat modeling methods (e.g., STRIDE) are insufficient to assess emerging GenAI-specific risks. In this work, we present the first results from an exploratory assessment of GenAI-aware threat modeling methods in a Small and Medium Enterprise (SME) setting. For this, we conducted a rapid literature review to select relevant techniques and systematically applied three shortlisted methods to an industrial case study involving a GenAI-augmented system. The results highlight differences in the threats identified by each technique and reveal limited support for certain GenAI-specific risk categories, particularly those related to software supply chains and human-centered security issues. We further report practitioners' perceptions of the usability and integration of these methods in SME development workflows, including their perceived effort and adoption challenges.
Cite as
Nicolás E. Díaz Ferreyra, Manish Mahesh Kumar, Nohemí Villarreal, Pankaj Pantel, Immo Brueggemann, and Riccardo Scandariato. Emerging Challenges in Threat Modeling for GenAI-Augmented Systems: A View from the Trenches. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 57:1-57:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{diazferreyra_et_al:LIPIcs.ESEM.2026.57,
author = {D{\'\i}az Ferreyra, Nicol\'{a}s E. and Kumar, Manish Mahesh and Villarreal, Nohem{\'\i} and Pantel, Pankaj and Brueggemann, Immo and Scandariato, Riccardo},
title = {{Emerging Challenges in Threat Modeling for GenAI-Augmented Systems: A View from the Trenches}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {57:1--57:15},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.57},
URN = {urn:nbn:de:0030-drops-280253},
doi = {10.4230/LIPIcs.ESEM.2026.57},
annote = {Keywords: Threat Modeling, Software Architectures, Generative AI, ML, Sec4AI}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Maximilian Alexander Amougou Mbida and Florian Angermeir
Abstract
Reproducibility in empirical software engineering relies on complete, accessible, and reusable research artifacts, yet artifact evaluation remains largely manual and difficult to scale. This emerging results paper explores an agentic approach for assessing replication package quality by translating open-science guidelines into machine-verifiable criteria. We consolidate 380 requirements from 34 sources into 51 reproducibility criteria, of which 31 are operationalized for automated artifact-based evaluation. Based on these criteria, we implement a multi-agent prototype that inspects replication packages and produces evidence-grounded improvement reports. A preliminary evaluation on five replication packages shows high inter-run consistency of 91.4% and 75.4% correctness, through micro-averaged agreement with a manual baseline. The agent performs best on structural criteria such as code, environment, and artifact availability, but struggles with qualitative or mixed-method studies. A pilot survey with seven software engineering researchers indicates well perceived usefulness and adoption potential, while revealing cognitive load in the human-in-the-loop planning step. Overall, within these small samples, the results indicate that agentic research artifact evaluation has the potential to support authors and reviewers by automating selected routine checks.
Cite as
Maximilian Alexander Amougou Mbida and Florian Angermeir. An Agentic Approach Towards Replication Package Quality Evaluation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 58:1-58:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{amougoumbida_et_al:LIPIcs.ESEM.2026.58,
author = {Amougou Mbida, Maximilian Alexander and Angermeir, Florian},
title = {{An Agentic Approach Towards Replication Package Quality Evaluation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {58:1--58:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.58},
URN = {urn:nbn:de:0030-drops-280266},
doi = {10.4230/LIPIcs.ESEM.2026.58},
annote = {Keywords: reproducibility, software engineering, research artifacts, agentic systems, multi-agent systems, evaluation framework, research infrastructure}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Antonino Coppola, Matteo Esposito, Rick Kazman, and Valentina Lenarduzzi
Abstract
Context. Generative AI coding agents are increasingly used to automate software maintenance and issue resolution. However, current evaluations mainly focus on test-passing behavior and provide limited insights into whether agents modify the same software entities selected by developers.
Aim. This study investigates the accuracy of issue localization performed by agents across different software granularities.
Method. We analyzed 2,441 issue-fixing commits from 10 large-scale Java projects and evaluated three agents, each combining the OpenCode harness with a different open-weight LLM. We compared agent-modified entities against human-implemented fixes at the package, class, and method levels using Accuracy, Precision, Recall, F1-score, and MCC.
Results. Agents partially identified the software entities requiring modification, but localization performance strongly depended on software granularity. Agents achieved the strongest results at the package level, while performance progressively degraded at the class and method levels. GLM-5 consistently achieved the strongest localization performance, although practical differences among agents remained limited. Our findings showed a limitation of agents implementing functionally plausible yet structurally different packages, class or methods leading to consistently negative MCC values worsening with finer granularity.
Conclusion. Current agents can often identify the general architectural region affected by an issue, but still struggle to precisely localize fine-grained implementation points. These findings highlight both the potential and the limitations of agents for issue localization and motivate larger-scale investigations on localization behavior, architectural impact, and long-term maintainability implications.
Cite as
Antonino Coppola, Matteo Esposito, Rick Kazman, and Valentina Lenarduzzi. On Coding Agent Issue Localization Accuracy - An Exploratory Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 59:1-59:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{coppola_et_al:LIPIcs.ESEM.2026.59,
author = {Coppola, Antonino and Esposito, Matteo and Kazman, Rick and Lenarduzzi, Valentina},
title = {{On Coding Agent Issue Localization Accuracy - An Exploratory Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {59:1--59:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.59},
URN = {urn:nbn:de:0030-drops-280271},
doi = {10.4230/LIPIcs.ESEM.2026.59},
annote = {Keywords: AI, Coding Agents, Empirical Study}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Agustín Olmedo, Jannik Laval, Christelle Urtado, and Sylvain Vauttier
Abstract
Large Language Models (LLMs) are increasingly embedded in software projects, yet little is known about how developers configure LLM parameters in the wild. Characterizing which parameters are defined, how many per project, which groups are co-defined, and what values are preferred can inform academics and tool builders about practices. This article mines open-source Python repositories on GitHub that integrate LLMs and uses AST-based static analysis to extract parameter assignments. Project-level definitions are analyzed to address prevalence (RQ1), configuration complexity and its structure (RQ2), and value distributions (RQ3). The final corpus includes 363 projects and gathers 7892 parameter definitions, with both the dataset and the code available for reproducibility. We observe that temperature is most frequently defined (90.36%), followed by max_tokens (56.75%), top_p (49.86%), and top_k (42.70%); penalty parameters are comparatively rare (RQ1). Projects typically define few parameters (mean ≈ 3), and this limited set expands incrementally around a stable core (temperature + max_tokens) (RQ2). Distributions suggest "defaults-in-practice": temperature ≈ 0.0, 0.7 and 1.0, top_p ≈ 0.9, top_k ≈ 0-50, and max_tokens at 2^k values (e.g., 256/512/1024) (RQ3). In conclusion, developers favor minimalist, sampling-centric configurations with convergence around specific ranges. These descriptive findings support clearer configuration reporting, offer practical baselines for tools and education, and motivate future work on causes, longitudinal evolution and task/domain stratification.
Cite as
Agustín Olmedo, Jannik Laval, Christelle Urtado, and Sylvain Vauttier. Parameterizing LLMs in Practice: An Empirical Study of LLMs Integrated into Software Systems. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 60:1-60:12, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{olmedo_et_al:LIPIcs.ESEM.2026.60,
author = {Olmedo, Agust{\'\i}n and Laval, Jannik and Urtado, Christelle and Vauttier, Sylvain},
title = {{Parameterizing LLMs in Practice: An Empirical Study of LLMs Integrated into Software Systems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {60:1--60:12},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.60},
URN = {urn:nbn:de:0030-drops-280280},
doi = {10.4230/LIPIcs.ESEM.2026.60},
annote = {Keywords: Software Engineering for AI, Large Language Model, Hyperparameter, LLM configuration, LLM API, Mining Software Repository}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Ronnie de Souza Santos, Italo Santos, Mauricio Rodrigues Lima, and Cleyton Magalhães
Abstract
Background. Interview-based studies are widely used in empirical software engineering to investigate human, organizational, and socio-technical phenomena, yet interview sample adequacy and saturation are reported inconsistently across the literature.
Aims. This paper investigates how interview sample adequacy and saturation are reported in empirical software engineering research.
Method. We analyzed 427 papers published between 2016 and 2025 across major software engineering venues, focusing on interview sample sizes, saturation discussions, and sample adequacy justifications.
Results. Preliminary findings indicate substantial variation in interview sample sizes, ranging from highly specialized small-sample studies to broader investigations involving large interview datasets. Studies involving 12 or fewer interviewees were common and frequently associated with specialized industrial contexts or constrained organizational access. However, the most common range was 13 to 24 interviewees, suggesting that moderate-sized samples represent the most common configuration in empirical software engineering research. Saturation and sample adequacy justifications were heterogeneous, with many studies relying on implicit or contextual reasoning rather than explicit methodological discussion.
Conclusions. Our findings provide empirical insights into methodological reporting practices in interview-based software engineering research and contribute to ongoing discussions on qualitative rigor and transparency.
Cite as
Ronnie de Souza Santos, Italo Santos, Mauricio Rodrigues Lima, and Cleyton Magalhães. How Many Interviews Are Enough in a Software Engineering Study? Preliminary Findings on Sample Size and Saturation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 61:1-61:12, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{desouzasantos_et_al:LIPIcs.ESEM.2026.61,
author = {de Souza Santos, Ronnie and Santos, Italo and Lima, Mauricio Rodrigues and Magalh\~{a}es, Cleyton},
title = {{How Many Interviews Are Enough in a Software Engineering Study? Preliminary Findings on Sample Size and Saturation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {61:1--61:12},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.61},
URN = {urn:nbn:de:0030-drops-280299},
doi = {10.4230/LIPIcs.ESEM.2026.61},
annote = {Keywords: qualitative research, interviews, publications}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Venla Liljas, Matteo Esposito, Valentina Lenarduzzi, and Davide Taibi
Abstract
Microservice architectures are increasingly becoming the standard design for cloud-native applications, yet their complex dependency structures make architectural understanding a challenge. Our research explores how Large Language Models (LLMs) can be leveraged to automatically reconstruct microservice architectures directly from source code, a task that requires both high-level architectural reasoning and fine-grained code understanding. We introduce a Minimum Viable Agent (MVA) as a single-agent that combines tool-assisted repository exploration with role-guided prompting to infer the services, inter-service connections, and exposed endpoints in Java microservice systems. We evaluate two state-of-the-art LLMs, GPT-4o and Llama4-16:17B, on 17 open-source microservice applications with 20 independent runs per system. GPT-4o achieves a mean F1 score of 0.604, outperforming Llama4 (F1 = 0.254). Both models show higher precision than recall, suggesting a tendency to under-represent architectural relationships rather than hallucinate them. Results also reveal substantial variation across applications and architectural element types. Analysis by architectural element indicates that service-to-service connections are the most challenging to recover for both models, while large performance differences emerge from endpoint reconstruction. Overall, our findings suggest that LLMs show promise for automated architecture recovery, but current capabilities remain insufficient for fully reliable reverse engineering. We discuss implications for integrating LLMs into architecture analysis workflows and outline directions for future research.
Cite as
Venla Liljas, Matteo Esposito, Valentina Lenarduzzi, and Davide Taibi. Can Agents Reconstruct Microservice Architecture?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 62:1-62:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{liljas_et_al:LIPIcs.ESEM.2026.62,
author = {Liljas, Venla and Esposito, Matteo and Lenarduzzi, Valentina and Taibi, Davide},
title = {{Can Agents Reconstruct Microservice Architecture?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {62:1--62:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.62},
URN = {urn:nbn:de:0030-drops-280302},
doi = {10.4230/LIPIcs.ESEM.2026.62},
annote = {Keywords: Large Language Models, Microservice Architecture, Software Architecture Reconstruction, Agentic AI}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Quanhe Wang, Xiangyun Zhan, Cheng Wen, Bin Yu, Ping Chen, Xingjian Han, and Shengchao Qin
Abstract
Large language models (LLMs) are increasingly used for code translation, yet existing evaluations often collapse results across language pairs, thereby obscuring the direction of transfer. We present an emerging empirical study that treats multilingual code translation as a directed transfer problem. Our study uses 100 programming problems, each with reference implementations in eight languages, yielding a complete directed graph of 56 ordered source-target translation directions. Five contemporary LLMs generate candidate translations, which are evaluated through execution in the target language. The results reveal substantial directional asymmetry. In model-averaged pass@5, reversing a language pair changes the translation success rate by as much as 27.4 percentage points. We also observe systematic differences between source and target roles: Python3 and Java achieve higher average performance when used as source languages, whereas JavaScript, Rust, and Golang achieve higher average performance when used as target languages. These findings show that aggregate scores and unordered language-pair averages can conceal practically important transfer behavior. Rather than establishing a definitive model ranking, this study provides preliminary evidence that code-translation evaluations should preserve ordered source-target results, distinguish source and target language roles, and explicitly report directional gaps. This direction-aware perspective provides a more informative basis for subsequent failure analysis and the evaluation of realistic cross-language migration scenarios.
Cite as
Quanhe Wang, Xiangyun Zhan, Cheng Wen, Bin Yu, Ping Chen, Xingjian Han, and Shengchao Qin. A Direction-Aware Study of LLM-Based Code Translation Across Eight Programming Languages. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 63:1-63:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{wang_et_al:LIPIcs.ESEM.2026.63,
author = {Wang, Quanhe and Zhan, Xiangyun and Wen, Cheng and Yu, Bin and Chen, Ping and Han, Xingjian and Qin, Shengchao},
title = {{A Direction-Aware Study of LLM-Based Code Translation Across Eight Programming Languages}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {63:1--63:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.63},
URN = {urn:nbn:de:0030-drops-280313},
doi = {10.4230/LIPIcs.ESEM.2026.63},
annote = {Keywords: Large language models, code translation, multilingual programming, empirical software engineering, execution-based evaluation}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Tanmay Singla and James C. Davis
Abstract
Modern software development relies on an increasingly doubtful premise: that the up-front implementation savings from adopting a dependency outweighs the maintenance costs. Two changes are reshaping the build-vs.-reuse calculus: software supply chain attacks have raised the cost of external reliance, while generative AI has lowered the cost of local implementation. In this vision paper, we discuss use-case-oriented regeneration as a new software sourcing paradigm that shifts the supply chain from external trust to local verification. We evaluate an agentic workflow that synthesizes only the specific slice of dependency functionality that a repository exercises. Our measurements across 180 repository-dependency pairs suggest that this approach is feasible: the replacements preserve 99.8% of repository-observed behavior across baseline validation checks and reduce the exported API surface by 93%. Software sourcing may evolve toward verifiable repository-specific code synthesis, especially when the required functionality is narrow, stable, and well tested.
Cite as
Tanmay Singla and James C. Davis. Software Supply Chains Are Dead: Use-Case-Oriented Regeneration. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 64:1-64:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{singla_et_al:LIPIcs.ESEM.2026.64,
author = {Singla, Tanmay and Davis, James C.},
title = {{Software Supply Chains Are Dead: Use-Case-Oriented Regeneration}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {64:1--64:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.64},
URN = {urn:nbn:de:0030-drops-280327},
doi = {10.4230/LIPIcs.ESEM.2026.64},
annote = {Keywords: Software supply chain, dependency reuse, code generation, AI coding agents}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Sabahat Younas, Márcia Moraes, and Fabio Santos
Abstract
Contributors to Open Source Software (OSS) projects are vital to maintaining the health of both communities and projects. However, the number of projects experiencing core contributors' disengagement has increased to the point that it risks the projects' survival. Finding new contributors to replace the workforce is challenging and time-intensive due to several factors, including a lack of precise knowledge about potential candidates' competences, which may need to be confirmed through interviews and exams. Previous studies provided indications of the contributor’s competences, although they lack depth in understanding competence levels, which can result in poor knowledge about the contributor’s capabilities. To address this gap, we assess contributors' competence by collecting code metrics related to source code from contributions. We also propose a competence model able to predict the competence level required to solve tasks. By properly assessing contributors' competences, we can identify and train candidates to replace core contributors. Our replication package, including code, data, and documentation, is available at https://doi.org/10.5281/zenodo.21605122
Cite as
Sabahat Younas, Márcia Moraes, and Fabio Santos. Towards Competence-Based Management for Open Source Software Projects. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 65:1-65:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{younas_et_al:LIPIcs.ESEM.2026.65,
author = {Younas, Sabahat and Moraes, M\'{a}rcia and Santos, Fabio},
title = {{Towards Competence-Based Management for Open Source Software Projects}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {65:1--65:15},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.65},
URN = {urn:nbn:de:0030-drops-280339},
doi = {10.4230/LIPIcs.ESEM.2026.65},
annote = {Keywords: Contributor competence, MSR, project sustainability, AST, OSS, disengagement mitigation}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Pasan Peiris, Matthias Galster, Antonija Mitrovic, Sanna Malinen, Raul Vincent Lumapas, and Jay Holland
Abstract
Soft skills are essential for software engineers. For example, effective communication and collaboration are critical for the design and delivery of software solutions. However, many organizations lack the resources to systematically support soft skills training. Although video-based training provides a flexible approach, it easily leads to shallow learning. To support engagement, active video watching integrates reflective activities, such as commenting on videos and reviewing peers’ comments. Gamifying these activities can further increase engagement. However, effective gamified training must carefully consider the target audience of the training. Therefore, this emerging results paper investigates the impact of gamification on the engagement, learning, and attitudes of software professionals participating in video-based soft skills training. We conducted an experiment involving 48 software professionals who completed a presentation skills training program. Our findings indicate that gamified video-based training did not significantly improve engagement or learning when compared to a non-gamified training. Nevertheless, participants generally expressed more positive attitudes toward gamified training. Based on these findings, we discuss several design considerations for technology-supported professional development for software professionals.
Cite as
Pasan Peiris, Matthias Galster, Antonija Mitrovic, Sanna Malinen, Raul Vincent Lumapas, and Jay Holland. Does Gamification Improve Video-Based Soft Skills Training of Software Engineering Professionals?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 66:1-66:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{peiris_et_al:LIPIcs.ESEM.2026.66,
author = {Peiris, Pasan and Galster, Matthias and Mitrovic, Antonija and Malinen, Sanna and Lumapas, Raul Vincent and Holland, Jay},
title = {{Does Gamification Improve Video-Based Soft Skills Training of Software Engineering Professionals?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {66:1--66:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.66},
URN = {urn:nbn:de:0030-drops-280344},
doi = {10.4230/LIPIcs.ESEM.2026.66},
annote = {Keywords: Software engineering, Soft Skills, Gamification, Video-based training, Empathy, Communication Skills, Professional development}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Muhammad Hamza, Dominik Siemon, and Wardah Naeem Awan
Abstract
AI coding agents such as GitHub Copilot and Claude Code are increasingly contributing pull requests (PRs) autonomously across large-scale software repositories. While prior research has primarily investigated the quality and acceptance of agent-generated artifacts, less is known about how sustained exposure to AI agents influences developers’ own practices. Drawing on automation complacency theory, we investigate whether agent adoption is associated with changes in coding effort, testing discipline, and code review behavior. We conduct a longitudinal study of 669 developers across 103 repositories using the AIDev dataset and analyze behavioral changes before and after agent adoption across 11 metrics. To distinguish agent-associated effects from broader ecosystem trends, we complement within-subject analyses with a difference-in-differences (DiD) design using 228 control developers from repositories without agent activity. We find one large agent-associated effect: a 76.8% increase in PR description length (r = 0.561), alongside consistent but smaller improvements in testing discipline across all three testing measures. Coding effort remains stable. Although review behavior changes noticeably - reviews become faster, less detailed, and more permissive - these patterns are also observed in control repositories and therefore appear to reflect broader ecosystem trends rather than agent adoption. Overall, we find no evidence that developers reduce productive effort when working with AI agents. Instead, developers increase documentation and testing activity, while apparent declines in review thoroughness are better explained by wider changes in software development practice. These findings not only clarify how developers adapt to AI coding agents but also demonstrate how comparative causal designs help distinguish agent-associated behavioral changes from broader ecosystem trends.
Cite as
Muhammad Hamza, Dominik Siemon, and Wardah Naeem Awan. Does Working with AI Agents Change How Developers Code, Test, and Review?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 67:1-67:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{hamza_et_al:LIPIcs.ESEM.2026.67,
author = {Hamza, Muhammad and Siemon, Dominik and Awan, Wardah Naeem},
title = {{Does Working with AI Agents Change How Developers Code, Test, and Review?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {67:1--67:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.67},
URN = {urn:nbn:de:0030-drops-280351},
doi = {10.4230/LIPIcs.ESEM.2026.67},
annote = {Keywords: AI Coding Agents, Automation Complacency, Developer Behavior, Pull Requests, Difference-in-Differences, Empirical Software Engineering, Longitudinal Study}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Minase Mekete Mengistu, Juri Di Rocco, Phuong T. Nguyen, and Davide Di Ruscio
Abstract
Background. Infrastructure-as-Code (IaC) scanners detect cloud misconfigurations in Terraform and other IaC languages before deployment, but repairing the flagged configurations remains largely manual. Recent Large Language Model (LLM)-based repair approaches can repair some findings, but may hallucinate unsupported constructs or suppress warnings without fixing the issue.
Aims. We study whether tool grounding can improve LLM-based Terraform repair, and when a finding should be escalated because the required deployment-specific context is not available.
Method. We present TerraRepair, a prototype of a tool-grounded LLM agent for Terraform repair with structured escalation. TerraRepair retrieves dependency context from Terraform references, consults the installed provider schema, and re-runs the scanner before returning a candidate repair. When the required context is absent, TerraRepair escalates instead of fabricating a plausible fix.
Results. We evaluate our tool on two vulnerable-by-design Terraform repositories using two IaC security scanners, Checkov and Trivy, across AWS, Azure, and GCP. On the combined AWS benchmark, TerraRepair improves scanner-verified fix rates from 26.6% to 78.4% on Checkov and from 44.8% to 72.4% on Trivy, compared with a controlled one-shot baseline. It also reduces the baseline’s 44.8-73.6 percentage point (pp) claimed-vs-verified repair gap to under 5 pp. In a sampled semantic audit covering AWS only, 78.9% of TerraRepair’s scanner-verified AWS repairs are labeled as correct under a majority-vote protocol with two LLM judges and one author.
Conclusions. These emerging results show that tool grounding can substantially improve scanner-verified LLM-based IaC repair on the studied benchmarks, while missing deployment-specific context remains the main knowledge boundary for full autonomy.
Cite as
Minase Mekete Mengistu, Juri Di Rocco, Phuong T. Nguyen, and Davide Di Ruscio. TerraRepair: A Tool-Grounded LLM Agent for Infrastructure-As-Code Repair. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 68:1-68:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{mengistu_et_al:LIPIcs.ESEM.2026.68,
author = {Mengistu, Minase Mekete and Di Rocco, Juri and Nguyen, Phuong T. and Di Ruscio, Davide},
title = {{TerraRepair: A Tool-Grounded LLM Agent for Infrastructure-As-Code Repair}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {68:1--68:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.68},
URN = {urn:nbn:de:0030-drops-280363},
doi = {10.4230/LIPIcs.ESEM.2026.68},
annote = {Keywords: Infrastructure as Code, automated repair, large language models, cloud security}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Darja Smite, Viggo Tellefsen Wivestad, Gabrielle O'Brien, Italo Santos, Giuseppe Destefanis, Maria Teresa Baldassarre, Mircea Lungu, Elda Paja, and Ronnie de Souza Santos
Abstract
As Artificial Intelligence (AI) assistants and agents become increasingly integrated into software engineering work, concerns about the economic and environmental costs of AI usage continue to grow. Current discussions of computational efficiency focus primarily on model architectures, hardware optimization, and inference cost reduction. In this vision paper, we argue that computational efficiency should also be understood as a socio-technical phenomenon shaped by patterns of human-AI interaction within organizational settings. Motivated by ongoing dialogues with industry partners and informed by the Theory of Planned Behavior and Social Cognitive Theory, we introduce the concept of computational behavior to describe patterns of AI-assisted work that influence how computational resources are consumed, reused, coordinated, and amplified across software engineering activities. Building on this perspective, we propose a multilevel conceptual framework linking individual computational behaviors to emergent collective computational outcomes. We argue that individually rational AI usage behaviors may accumulate into either computational waste or collective computational efficiency depending on how interactions are shared, reused, and aligned across teams and workflows. While fragmented AI use may generate duplicated prompting, redundant generation, and coordination overhead, coordinated collective practices, such as reusable prompts, shared contextualization, and workflow integration may improve collective computational efficiency. We conclude by outlining a research agenda for studying computational coordination, including empirical and organizational AI governance in AI-assisted software engineering.
Cite as
Darja Smite, Viggo Tellefsen Wivestad, Gabrielle O'Brien, Italo Santos, Giuseppe Destefanis, Maria Teresa Baldassarre, Mircea Lungu, Elda Paja, and Ronnie de Souza Santos. Beyond Individual Prompting Efficiency: A Socio-Technical Perspective on Computational Efficiency in AI-Assisted Software Engineering. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 69:1-69:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{smite_et_al:LIPIcs.ESEM.2026.69,
author = {Smite, Darja and Wivestad, Viggo Tellefsen and O'Brien, Gabrielle and Santos, Italo and Destefanis, Giuseppe and Baldassarre, Maria Teresa and Lungu, Mircea and Paja, Elda and de Souza Santos, Ronnie},
title = {{Beyond Individual Prompting Efficiency: A Socio-Technical Perspective on Computational Efficiency in AI-Assisted Software Engineering}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {69:1--69:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.69},
URN = {urn:nbn:de:0030-drops-280377},
doi = {10.4230/LIPIcs.ESEM.2026.69},
annote = {Keywords: AI-assisted software engineering, Computational waste, Team practices, Mental models, Empirical software engineering, Vision, Research agenda}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Lam D. Dao, Vang T. Nguyen, Anh M. T. Bui, and Phuong T. Nguyen
Abstract
Background. Large language models (LLMs) have become increasingly capable of understanding and generating source code, leading to their widespread adoption in software engineering tasks such as code completion, repair, and vulnerability detection. However, despite their strong empirical performance, the internal mechanisms through which LLMs recognize malicious or vulnerable code patterns remain poorly understood.
Aim. We investigated where the malware detection behavior is encoded inside LLMs Feed Forward Network (FFN) neurons and verified the attribution with causal interventions on the neurons identified. This aims to identify the most important neurons in detecting malicious code.
Methods. We applied mechanistic interpretability methods to locate the neurons being responsible for malware-detection behavior in three instruction-tuned LLMs: Llama3.1-8B-Instruct, Mistralv0.3-7B-Instruct, and Qwen2.5-7B-Instruct. Using 1,500 malicious and 1,500 benign PyPI packages from the PyPI Malregistry, we attribute the behavior to a set of neurons.
Results. The experimental results reveal that amplifying facilitating neurons for malware detection while suppressing inhibiting ones can boost accuracy, while the reverse collapses predictions toward a single class, although the magnitude and consistency is heavily model-dependent. We demonstrated that the guardrail detection mechanism varies across models, each represents its malware detection behavior differently within its FFN layers.
Conclusions. Probing the neurons associated with security-relevant knowledge helps us gain insights into how LLMs encode malicious programming concepts, identify potentially harmful memorized behaviors, paving the way toward more reliable defense mechanisms, such as neuron-level editing, selective unlearning, and security-aware alignment for code-focused LLMs.
Cite as
Lam D. Dao, Vang T. Nguyen, Anh M. T. Bui, and Phuong T. Nguyen. Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 70:1-70:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{dao_et_al:LIPIcs.ESEM.2026.70,
author = {Dao, Lam D. and Nguyen, Vang T. and Bui, Anh M. T. and Nguyen, Phuong T.},
title = {{Which Neurons Detect Malicious Code? A Probing Study of LLM Security Knowledge}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {70:1--70:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.70},
URN = {urn:nbn:de:0030-drops-280385},
doi = {10.4230/LIPIcs.ESEM.2026.70},
annote = {Keywords: Malicious code, Unlearning methods, LLMs}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Shalini Chakraborty and Jan-Philipp Steghöfer
Abstract
AI coding assistants are fundamentally reshaping software development by shifting developers' effort from writing code toward specifying intent through natural language prompts. In emerging chat-based development practices such as vibe coding, prompts mediate the transformation of human intent into executable software. While Requirements Engineering (RE) emphasizes capturing, validating, and evolving requirements, current prompting practices remain informal and ad hoc. In this vision paper, we argue that prompts represent lightweight, evolving requirements artifacts that combine expressions of user needs with varying degrees of solution guidance. We use an existing conceptual model that decomposes prompts into three interrelated dimensions: Functionality and Quality (capturing intended system requirements), General Solutions (capturing architectural strategies and technology choices), and Specific Solutions (capturing implementation-level constraints and directives). Building on this conceptualization, we formulate four research hypotheses concerning (i) the evolution of prompts over time, (ii) the influence of user characteristics on prompt evolution, (iii) the relationship between prompt content and requirements validation and verification activities, and (iv) the impact of prompt characteristics on requirements and resulting software quality. We envision an empirical research agenda combining real-world AI-assisted development data, corpus analysis, and controlled experimentation to investigate these hypotheses and derive evidence-based practices for requirements-aware prompt engineering. By reframing prompts through the lens of RE, we position prompting not merely as an interaction mechanism with AI systems, but as a central software engineering concern requiring systematic study.
Cite as
Shalini Chakraborty and Jan-Philipp Steghöfer. Prompts Blend Requirements and Solutions: From Intent to Implementation. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 71:1-71:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{chakraborty_et_al:LIPIcs.ESEM.2026.71,
author = {Chakraborty, Shalini and Stegh\"{o}fer, Jan-Philipp},
title = {{Prompts Blend Requirements and Solutions: From Intent to Implementation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {71:1--71:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.71},
URN = {urn:nbn:de:0030-drops-280393},
doi = {10.4230/LIPIcs.ESEM.2026.71},
annote = {Keywords: Requirements Engineering, Vibe Coding, Prompts, AI}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Yutan Huang, Chetan Arora, Tanjila Kanij, Anuradha Madugalla, and John Grundy
Abstract
Healthcare AI software systems raise ethical concerns that are often discussed as high-level principles but are difficult to translate into concrete software engineering requirements. This emerging-results paper reports preliminary findings from a multi-stakeholder interview study on how practitioners experience and prioritise ethical concerns in healthcare AI software systems. We conducted 17 semi-structured interviews with healthcare professionals (n=5), healthcare researchers (n=6), and software engineers (n=6). Using thematic synthesis, we identified three AI use contexts and four recurring challenge categories: model reliability and quality, data quality and processing, human oversight and skills gaps, and institutional and regulatory constraints. Privacy and safety emerged as baseline concerns across roles, while bias, accountability, and transparency varied depending on practitioners' responsibilities and use contexts. Based on these preliminary findings, we propose role-specific patterns as lightweight requirements artefacts to translate practitioner priorities into actionable software engineering practices across requirements elicitation, system design, testing, and auditing. We present these patterns as an initial artefact for community feedback and as a first step toward a broader approach to ethical requirements engineering (RE) for healthcare AI software systems. Future work will validate and refine the patterns using a larger, more balanced sample, evaluate their usefulness through practitioner workshops, and develop a reusable catalogue with guidance on traceability, verification, and auditing. Long-term, our goal is to support healthcare AI teams in moving from abstract ethical principles to testable requirements that can be integrated into existing development and governance workflows.
Cite as
Yutan Huang, Chetan Arora, Tanjila Kanij, Anuradha Madugalla, and John Grundy. Ethical Requirements for Healthcare AI Software Systems: Emerging Results from a Multi-Stakeholder Interview Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 72:1-72:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{huang_et_al:LIPIcs.ESEM.2026.72,
author = {Huang, Yutan and Arora, Chetan and Kanij, Tanjila and Madugalla, Anuradha and Grundy, John},
title = {{Ethical Requirements for Healthcare AI Software Systems: Emerging Results from a Multi-Stakeholder Interview Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {72:1--72:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.72},
URN = {urn:nbn:de:0030-drops-280406},
doi = {10.4230/LIPIcs.ESEM.2026.72},
annote = {Keywords: Healthcare AI software systems, ethical requirements, requirements engineering, empirical software engineering, thematic synthesis, practitioner interviews}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Mary Sánchez-Gordón, Ricardo Colomo-Palacios, and Rahul Mohanani
Abstract
The software testing process is inherently adversarial, creating conditions for frequent and pervasive developer–tester conflicts. Although forgiveness has been studied as a tool for resolving conflicts and as a coping mechanism in the workplace context that promotes well-being and productivity, its role in software testing remains unexplored. This paper posits that cultivating forgiveness can help practitioners manage their emotional responses to conflict and therefore deserves further investigation in the context of software testing.
Cite as
Mary Sánchez-Gordón, Ricardo Colomo-Palacios, and Rahul Mohanani. When Work Relations Matter: A Vision of Forgiveness in Software Testing. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 73:1-73:6, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{sanchezgordon_et_al:LIPIcs.ESEM.2026.73,
author = {S\'{a}nchez-Gord\'{o}n, Mary and Colomo-Palacios, Ricardo and Mohanani, Rahul},
title = {{When Work Relations Matter: A Vision of Forgiveness in Software Testing}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {73:1--73:6},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.73},
URN = {urn:nbn:de:0030-drops-280415},
doi = {10.4230/LIPIcs.ESEM.2026.73},
annote = {Keywords: Forgiveness, Emotional regulation, Conflicts, Software testing}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Chukwudi Uwasomba, Helen Sharp, Tamara Lopez, Michel Wermelinger, Advait Deshpande, and Peggy Gregory
Abstract
Large-Scale Agile Frameworks (LSAFs) are widely used to coordinate agile transformation beyond individual teams. They are often associated with organisational adaptability, yet whether their formal guidance encodes the capabilities associated with socio-technical resilience (STR) is unclear. Drawing on Hollnagel’s resilience potentials, this vision paper presents observations from an initial investigation of how resilience-relevant capabilities are represented in formal LSAF guidance. These observations indicate a potential resilience blind spot in LSAFs. The paper identifies opportunities for further empirical work on LSAFs and resilience, and on how organisations supplement LSAFs with resilience-oriented practices.
Cite as
Chukwudi Uwasomba, Helen Sharp, Tamara Lopez, Michel Wermelinger, Advait Deshpande, and Peggy Gregory. Do Large-Scale Agile Frameworks Encode Socio-Technical Resilience?. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 74:1-74:6, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{uwasomba_et_al:LIPIcs.ESEM.2026.74,
author = {Uwasomba, Chukwudi and Sharp, Helen and Lopez, Tamara and Wermelinger, Michel and Deshpande, Advait and Gregory, Peggy},
title = {{Do Large-Scale Agile Frameworks Encode Socio-Technical Resilience?}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {74:1--74:6},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.74},
URN = {urn:nbn:de:0030-drops-280422},
doi = {10.4230/LIPIcs.ESEM.2026.74},
annote = {Keywords: Large-scale agile frameworks, Socio-technical resilience, Resilience engineering}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Niruthiha Selvanayagam and Taher A. Ghaleb
Abstract
AI coding agents are increasingly integrated into software development workflows, operating on both sides of the pull request (PR) process: AI authoring agents, which create or modify PRs, and AI reviewers, which evaluate them. This creates a closed loop where one AI coding agent reviews contributions of another AI coding agent. In this paper, we construct an AI-to-AI code review dataset by linking AI-authored pull requests with AI-attributed review events from CodAGE, a public dataset of coding agent–generated GitHub events. Our dataset contains νm{248641} PRs across 12 AI coding agents with at least one AI review, including νm{45269} reviewed by a different AI product and νm{208145} by the same product. We observe that cross-product AI-to-AI code review occurs in only about 1.6% of identified agent-authored PRs but is substantial in absolute terms: 45k PRs written by one identifiable AI product and reviewed by another. This activity grows by more than two orders of magnitude from 2025-Q1 to 2025-Q3. We measure reviewer behavior using CodeRabbit comment categories, per-PR comment volume, and time to first review, and find that it varies across author–reviewer pairs. For example, Claude-Code PRs receive more refactor comments from CodeRabbit than Copilot PRs (35.0% vs. 10.5%), a difference that may stem from PRs themselves rather than the reviewer. For three of four dual-role reviewers, mean comments per PR were 58-65% higher in the same-product group, though effect sizes were small or negligible and the difference was concentrated in the upper tail. Median time from PR creation to first AI review was 1.2 minutes for cross-product pairs and 4.7 minutes for same-product pairs, reflecting which reviewer bots are in each group rather than the product pairing. Overall, our large-scale characterization shows that closed-loop AI-to-AI code review is on the rise but remains a minority phenomenon, with review output varying across authoring-agent groups and author-reviewer configurations.
Cite as
Niruthiha Selvanayagam and Taher A. Ghaleb. AI-to-AI Code Reviews of GitHub Pull Requests. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 75:1-75:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{selvanayagam_et_al:LIPIcs.ESEM.2026.75,
author = {Selvanayagam, Niruthiha and Ghaleb, Taher A.},
title = {{AI-to-AI Code Reviews of GitHub Pull Requests}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {75:1--75:15},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.75},
URN = {urn:nbn:de:0030-drops-280430},
doi = {10.4230/LIPIcs.ESEM.2026.75},
annote = {Keywords: AI coding agents, AI code review, closed-loop AI, Pull requests, Mining software repositories, GitHub events}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Ryan Dang, Noshin Tahsin, and Thomas Zimmermann
Abstract
Background. Threats to validity (TTV) are a core component of empirical research, yet software engineering studies fail to address them adequately. Published studies often omit validity discussions entirely, report generic concerns disconnected from the study’s actual methodology, or treat validity analysis as a procedural formality rather than a purposeful part of the research process.
Aim. This study investigates whether large language models (LLMs) can generate threats to validity that are meaningful, study-specific, and novel, i.e., not reported by the original authors.
Method. We conducted an empirical experiment using Gemini 2.5 Flash across 375 papers from ICSE 2025, withholding each paper’s TTV section prior to model input, ensuring the model reasoned from the study’s content rather than reproducing author-reported concerns. Generated threats were evaluated against a four-dimension rubric covering relevance, specificity, clarity, and mitigation quality, and compared against the original threats to determine novelty. All threats were then grouped into a taxonomy of seven categories to characterize the distribution of generated validity concerns.
Results. The experiment produced 2,673 threats across the dataset, of which 97% were rated as relevant, specific, and clearly articulated, and 92% received high ratings for mitigation quality. Furthermore, 61% of the generated threats were absent from the original papers, suggesting that the model identified validity concerns beyond those reported by the original authors.
Conclusions. LLMs can generate technically sound and previously unreported threats to validity at scale, demonstrating their capacity to support researchers in identifying overlooked methodological risks and strengthening the completeness of validity discussions, while augmenting rather than replacing researcher judgment in the process.
Cite as
Ryan Dang, Noshin Tahsin, and Thomas Zimmermann. Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 76:1-76:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{dang_et_al:LIPIcs.ESEM.2026.76,
author = {Dang, Ryan and Tahsin, Noshin and Zimmermann, Thomas},
title = {{Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {76:1--76:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.76},
URN = {urn:nbn:de:0030-drops-280443},
doi = {10.4230/LIPIcs.ESEM.2026.76},
annote = {Keywords: Threats to Validity, Empirical Research, Large Language Models, LLMs, Software Engineering Research, Automated Analysis}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Sahand Moslemi Yengejeh and Anil Koyuncu
Abstract
Automated program repair (APR) tools generate candidate patches and accept those that pass a given test suite. A test-passing patch, called a plausible patch, is not necessarily correct: it may overfit the test suite without fixing the underlying bug. Automated patch correctness assessment (APCA) tools were developed to detect such overfitting patches, but they have been evaluated only as post-hoc classifiers on static datasets, never inside the repair loop they are intended to support. We investigate whether placing an APCA tool inside the generate-and-validate loop of an APR system, as a second validation gate, affects the APCA’s classification effectiveness, the APR tool’s repair capability, and the search cost it incurs. In this design, a candidate that passes all tests but is classified as overfitting by the APCA tool is discarded, and the repair tool continues generating candidates. We study it with three APR tools (TBar, ARJA, ChatRepair) crossed with three APCA tools (ODS, Quatrain, LLM4PatchCorrect), evaluated on 549 single-method bugs of Defects4J v3 with manual correctness assessment of 1,328 patches drawn from the in-loop candidate streams, and paired non-parametric tests (McNemar, Wilcoxon) per (APR, APCA) cell. Average in-loop F₁ on the correct class sits between 0.43 and 0.66 across the nine cells, with classification harder in-loop than post-hoc on every tool. The in-loop gate raises the share of correct outputs in 14 of 18 cells, but cuts the absolute correct-fix count whenever the APR baseline’s first plausible patch is already correct at a high rate. The in-loop introduces a search-cost overhead that is statistically significant in 14 of 18 cells (Wilcoxon p < .05), with a geometric-mean per-bug multiplier of up to 17.46×. Placing an APCA inside the repair loop raises the output correctness rate in most (APR, APCA) cells, but at a search-cost overhead and, where the APR baseline already produces correct first plausibles at a high rate, a reduction in absolute correct fixes. The findings indicate that APCA tools should be trained and evaluated on the in-loop patch distribution they are deployed against, rather than on the curated, terminal patches of static post-hoc datasets.
Cite as
Sahand Moslemi Yengejeh and Anil Koyuncu. APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 77:1-77:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{yengejeh_et_al:LIPIcs.ESEM.2026.77,
author = {Yengejeh, Sahand Moslemi and Koyuncu, Anil},
title = {{APCA in the Loop: An Empirical Study of In-Loop Patch Correctness Assessment for APR}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {77:1--77:15},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.77},
URN = {urn:nbn:de:0030-drops-280458},
doi = {10.4230/LIPIcs.ESEM.2026.77},
annote = {Keywords: automated program repair, patch correctness assessment, patch overfitting}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Lanaya A. Carbonell and Anıl Koyuncu
Abstract
Adding vulnerability context to a prompt is known to improve LLM repair quality, but the studies that establish this vary two things at once: the domain knowledge injected, such as a CWE label or a fix pattern, and the form that knowledge takes in the prompt. Prompt-engineering research treats format as a variable in its own right, yet in automated vulnerability repair it has never been isolated under a controlled ablation, leaving open whether the wording of guidance matters once its content is fixed. This paper presents CWEFT, a controlled ablation that holds injected repair knowledge constant across a free-text and a typed-schema rendering while decomposing the separate contributions of naming a vulnerability and supplying fix guidance. We construct four prompt conditions that differ by exactly one factor each, grounding all injected knowledge in APR4Vul’s empirically mined fix patterns. Using this design, we generate 504 patches across 42 Java vulnerabilities from Vul4J spanning 15 CWE classes on three frontier models - Claude Sonnet 4.5, GPT-5, and Gemini 2.5 Pro - and score each patch on a Proof-of-Vulnerability-aware hierarchy that separates code that compiles from code that closes the vulnerability and from code that does so without regression. Our emerging results invert a common assumption: naming the CWE class alone achieves the highest correct-fix rate at 19.8%, ahead of the 12.7% unenriched baseline (p = 0.049, exact McNemar), while both richer conditions that add explicit repair guidance fall back to 16.7%. Crucially, the free-text and typed-schema conditions are indistinguishable at 16.7% with perfectly symmetric paired discordance, and with only eight discordant pairs the design cannot rule out a smaller format effect. Across all conditions, repair rate tracks how concrete the vulnerability class is, falling to 3.1% for improper input validation. These initial findings suggest that for frontier models on common vulnerability classes, what a prompt identifies matters more than how elaborately that information is formatted - a direction that warrants higher-powered replication before it is treated as a design principle.
Cite as
Lanaya A. Carbonell and Anıl Koyuncu. CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 78:1-78:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{carbonell_et_al:LIPIcs.ESEM.2026.78,
author = {Carbonell, Lanaya A. and Koyuncu, An{\i}l},
title = {{CWEFT: CWE-Aware Evaluation of Free-Text vs. Typed Prompts}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {78:1--78:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.78},
URN = {urn:nbn:de:0030-drops-280469},
doi = {10.4230/LIPIcs.ESEM.2026.78},
annote = {Keywords: Automated vulnerability repair, Automated program repair, Large language models, Prompt engineering, Structured prompting, Software security, Vul4J}
}
Document
Emerging Results, Vision & Reflection Track Paper
Authors:
Haibo Yu and Jianjun Zhao
Abstract
Quantum software is often described using circuit-level metrics, including the number of qubits, the number of gates, circuit depth, the number of two-qubit gates, and the number of measurement operations. These metrics are useful for estimating program size and execution cost, but they do not directly describe whether and to what extent qubits become structurally related through possible entanglement during program execution. This paper argues that entanglement dependence provides a basis for measuring a quantum-specific form of software complexity. The proposed metrics capture possible entanglement-related structural dependence, not physical entanglement observed in a particular execution. Based on this notion, we define a focused set of six metrics that characterize the amount, density, local connectivity, global scope, program span, and modular crossing of entanglement-related dependencies. We also analyze these metrics using Weyuker’s properties for evaluating software complexity measures. The paper presents the metrics as a measurement framework and research direction for empirical quantum software engineering. It discusses how future studies can validate the metrics and examine their usefulness for testing, debugging, maintenance, and program comprehension of quantum software.
Cite as
Haibo Yu and Jianjun Zhao. Toward Entanglement Dependence Metrics for Quantum Software Complexity. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 79:1-79:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{yu_et_al:LIPIcs.ESEM.2026.79,
author = {Yu, Haibo and Zhao, Jianjun},
title = {{Toward Entanglement Dependence Metrics for Quantum Software Complexity}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {79:1--79:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.79},
URN = {urn:nbn:de:0030-drops-280472},
doi = {10.4230/LIPIcs.ESEM.2026.79},
annote = {Keywords: Quantum software engineering, quantum program analysis, entanglement dependence, software metrics, software complexity}
}
Document
Software Engineering in Practice Track Paper
Authors:
Yiyang Peng, Gang Lv, Yifeng Gou, Qizhao Wang, Bo Zhang, Ming Zhang, Qing Xu, Xiaosong Ding, Yu Shen, Yusang Xiong, Xinyao Xiao, and Xinyu Liu
Abstract
Industrial LLM compliance-auditing systems are maintained through both expert policy and executable support chains. After deployment, changing standards, user appeals, expert reviews, and report variants expose failures in rule interpretation, parsing, evidence binding, context assembly, output contracts, and regression checks. We propose CoEvolve, a feedback-driven policy-harness maintenance workflow in which an LLM maintenance agent proposes coordinated updates, an independent verifier checks target repair, regression, stability, and semantic risks, and experts confirm releasable changes. We report a five-month industrial experience in a Chinese compliance-auditing service, covering 33 production feedback instances, 128 audited systems, 384 reports, and 24,192 checklist items. The evaluation compares four frozen workflow states formed by different maintenance organizations; it is not a controlled comparison of methods on identical maintenance inputs. On a 979-item monthly holdout benchmark, the CoEvolve replay state reached 93.5% accuracy, 4.1 percentage points above prompt/SOP maintenance at 89.4%. Manual co-maintenance reached 93.6% (916/979), while CoEvolve reached 915/979; their Wilson intervals overlap and the five-system effective sample is small. Because the replay input included visible historical manual artifacts, this result shows that CoEvolve reconstructed a comparable release state from those artifacts, not that it substituted for manual co-maintenance. Across replay tasks, CoEvolve externalized 4.97 candidate attempts per feedback instance on average. The experience supports governed policy-harness maintenance without retraining and highlights verifier-gate predicates and blocked-candidate records as transferable release-governance mechanisms.
Cite as
Yiyang Peng, Gang Lv, Yifeng Gou, Qizhao Wang, Bo Zhang, Ming Zhang, Qing Xu, Xiaosong Ding, Yu Shen, Yusang Xiong, Xinyao Xiao, and Xinyu Liu. CoEvolve: Feedback-Driven Policy-Harness Maintenance for Industrial LLM Auditing Systems. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 80:1-80:18, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{peng_et_al:LIPIcs.ESEM.2026.80,
author = {Peng, Yiyang and Lv, Gang and Gou, Yifeng and Wang, Qizhao and Zhang, Bo and Zhang, Ming and Xu, Qing and Ding, Xiaosong and Shen, Yu and Xiong, Yusang and Xiao, Xinyao and Liu, Xinyu},
title = {{CoEvolve: Feedback-Driven Policy-Harness Maintenance for Industrial LLM Auditing Systems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {80:1--80:18},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.80},
URN = {urn:nbn:de:0030-drops-280480},
doi = {10.4230/LIPIcs.ESEM.2026.80},
annote = {Keywords: LLM auditing, policy-harness maintenance, production feedback, verifier gates, software engineering in practice}
}
Document
Software Engineering in Practice Track Paper
Authors:
Simin Sun, Peter Nemeth, Staffan Johansson, Theo Wiik, David Friberg, Farnaz Fotrousi, and Miroslaw Staron
Abstract
Software quality control in industry often relies on expensive analyses that run continuously in the Continuous Integration (CI) pipeline. Those analyses, such as static analysis for rule-based compliance, can provide important quality signals and often serve as strict gating points in the CI pipeline to ensure strict quality standards. However, strict continuous gating can slow the pipeline especially in the shared development environment. This study investigates the adoption of different software quality control strategies in safety-critical domain. By changing when the analyses take place and how the information can be delivered, we aim to identify a strategy that improves CI throughput while maintaining high standards of software quality. We conducted action research at a large safety-critical software organization for autonomous driving. Through this process, we investigated existing quality strategies and designed a new strategy that combine nightly analysis with feedback mechanism. We evaluate this strategy using both longitudinal data on static analysis issues and development activities, together with interview data with practitioners. The evaluation shows that nightly CI runs together with feedback mechanism helped keep selected quality issues under control during daily development and reduced reliance on clean-up in later stages. It can also improve the teams' follow-up responsibility and development efficiency. This study contributes an empirically grounded quality control strategy that mitigates the tradeoff between strict quality gate and CI pipeline throughput in safety-critical software development. It also identifies the organizational preconditions, implementation instructions and possible social effects for applying similar mechanisms in other environments that rely on expensive quality analyses.
Cite as
Simin Sun, Peter Nemeth, Staffan Johansson, Theo Wiik, David Friberg, Farnaz Fotrousi, and Miroslaw Staron. A Nightly Feedback-Driven Strategy for Quality Control in Safety-Critical Software Development. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 81:1-81:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{sun_et_al:LIPIcs.ESEM.2026.81,
author = {Sun, Simin and Nemeth, Peter and Johansson, Staffan and Wiik, Theo and Friberg, David and Fotrousi, Farnaz and Staron, Miroslaw},
title = {{A Nightly Feedback-Driven Strategy for Quality Control in Safety-Critical Software Development}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {81:1--81:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.81},
URN = {urn:nbn:de:0030-drops-280497},
doi = {10.4230/LIPIcs.ESEM.2026.81},
annote = {Keywords: Safety-critical Systems, Software Quality Control, Continuous Integration/Continuous Deployment (CI/CD)}
}
Document
Software Engineering in Practice Track Paper
Authors:
Vinay Kabadi, Abhiram Wuntakal, Naresh Gutha, Lahiri Bellarykar, Xuan Bach D. Le, Patanamon Thongtanunam, and Christoph Treude
Abstract
Business Requirements often fail to fully capture the complete set of system requirements they intend to describe, resulting in under-specification and measurable gaps between documented requirements and production-ready implementations. While experienced business analysts attempt to mitigate these gaps using domain expertise, the resulting artifacts often remain incomplete.
This study uses Cross-Pollination, a domain-agnostic methodology for transforming incomplete natural language Business Requirements into comprehensive, production-grade functional user stories. To evaluate the completeness of LLM-generated user stories, the study systematically benchmarks them against realistic production artifacts. It further investigates the capability of Large Language Models (LLMs) to bridge the coverage gap between human-authored backlogs and production systems. It identifies the requirements that are overlooked during human-driven elicitation but are recovered through an LLM-assisted approach.
The cross-pollination technique is a structured multi-phase methodology supported by a gap analysis taxonomy and a domain complexity calibration mechanism to establish a minimum baseline for production readiness. The methodology is applied across real-world domains, and a comparative evaluation of analyst-authored, LLM-generated, and production-derived user stories is conducted. Empirical results demonstrate that a substantial proportion of requirement gaps can be systematically identified and addressed through this approach.
The findings indicate that Cross-Pollination provides a reproducible and scalable methodology for translating natural-language requirements into more complete and production-aligned user stories. The results demonstrate that while analysts produce precise requirements, LLM-generated user stories achieve higher recall across all evaluated domains and are effective in uncovering implicit and non-obvious requirements. These findings indicate that an integrated approach, with LLM-assisted elicitation functioning as a complementary mechanism to human expertise, produces more comprehensive and reliable requirement backlogs than either method independently.
Cite as
Vinay Kabadi, Abhiram Wuntakal, Naresh Gutha, Lahiri Bellarykar, Xuan Bach D. Le, Patanamon Thongtanunam, and Christoph Treude. Can LLMs Identify Missing Elements in Requirements Elicitation? An Empirical Evaluation Using Production-Derived User Stories in an Industry Setting. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 82:1-82:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{kabadi_et_al:LIPIcs.ESEM.2026.82,
author = {Kabadi, Vinay and Wuntakal, Abhiram and Gutha, Naresh and Bellarykar, Lahiri and Bach D. Le, Xuan and Thongtanunam, Patanamon and Treude, Christoph},
title = {{Can LLMs Identify Missing Elements in Requirements Elicitation? An Empirical Evaluation Using Production-Derived User Stories in an Industry Setting}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {82:1--82:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.82},
URN = {urn:nbn:de:0030-drops-280500},
doi = {10.4230/LIPIcs.ESEM.2026.82},
annote = {Keywords: Software Engineering, Requirement Elicitation, Large Language Model (LLM), User Stories}
}
Document
Software Engineering in Practice Track Paper
Authors:
António Azevedo, Bruno Lima, and João Pascoal Faria
Abstract
Modern automotive infotainment systems sit at the center of the vehicle’s digital ecosystem, yet their validation still relies heavily on manual testing, a process that is time-consuming, expensive, and increasingly incompatible with the pace of agile release cycles and over-the-air software updates. Traditional scripted automation offers only a partial remedy, as it creates tight coupling between test logic and implementation details, producing brittle suites with high maintenance overhead. Existing LLM-driven testing frameworks predominantly target web or mobile applications, while employing single-agent or dual-agent architectures that overload one or two models with perception, planning, action selection, and validation simultaneously, making them prone to hallucination-style failures and unproductive exploration loops when faced with the complexity of automotive infotainment interfaces. In this paper, we present ARIA (Autonomous Real-time Infotainment Assessment), a multi-agent framework that leverages Large Language Models to autonomously execute end-to-end test scenarios on Android-based infotainment systems through visual interface interaction, orchestrating a closed-loop pipeline of four specialized agents per interaction step, complemented by a dedicated report-generation stage. From single-sentence natural-language scenario descriptions alone, each specifying a navigation path, an action to perform, and an expected outcome to verify, it autonomously executes the corresponding interactions on the infotainment system and produces structured reports, reproducible action scripts, and visual evidence for each step. ARIA was evaluated in an industrial setting on a physical test environment running the Android-based infotainment system of a car manufacturer, executing 30 scenarios spanning diverse system functionalities. Of the 30 scenarios, 28 (93.3%) completed the full multi-agent pipeline and produced a verdict, while 2 terminated prematurely with execution errors. Of the 28 completed scenarios, 20 (71.4%) matched the ground truth. ARIA detected all 5 known functional defects in the test setup, so no genuine fault was ever passed as working; the 8 false positives among completed scenarios are attributable to LLM navigation and image-interpretation limitations and to unsupported interaction gestures. These findings demonstrate that multi-agent LLM architectures can autonomously execute end-to-end infotainment test scenarios in an industrial setting, while also exposing the precision challenges that a low false-positive tolerance imposes. A single-agent baseline ablation on the same scenarios confirms the value of the multi-agent decomposition: on the first pass, before revisitation with a stronger model masks the difference, the single agent produces a substantially higher false positive rate (72.0% versus 52.6%) due to the conflation of navigational difficulty with system failure. We separately report first-pass and post-revisitation results and quantify token consumption, LLM-call counts, and monetary cost per scenario for both pipelines, and repeated execution of a representative subset of scenarios confirms that outcome stability correlates with scenario complexity, with fault detection remaining perfectly consistent across runs, indicating a path toward integrating visual test execution into continuous integration pipelines.
Cite as
António Azevedo, Bruno Lima, and João Pascoal Faria. ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 83:1-83:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{azevedo_et_al:LIPIcs.ESEM.2026.83,
author = {Azevedo, Ant\'{o}nio and Lima, Bruno and Faria, Jo\~{a}o Pascoal},
title = {{ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {83:1--83:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.83},
URN = {urn:nbn:de:0030-drops-280517},
doi = {10.4230/LIPIcs.ESEM.2026.83},
annote = {Keywords: Testing, AI, LLM, Agentic, Infotainment, Autonomous, UI}
}
Document
Software Engineering in Practice Track Paper
Authors:
Brenn Hill
Abstract
Context. Spec-driven development (SDD) tools claim that writing specifications before implementation reduces defects, prevents rework, and improves code quality. No vendor has published empirical evidence for these claims.
Method. We test five hypotheses derived from SDD vendor claims against 88,052 pull requests across 119 open-source repositories. What we measure is specification artifacts - overwhelmingly references to tracking issues and tickets - rather than SDD tool output, which no repository in the sample commits. We score 25,209 artifacts on the same quality dimensions SDD tools prescribe, measure rework from subsequent pull request activity, trace defects via the SZZ algorithm, and compare each developer’s spec'd PRs to their own unspec'd PRs. Twelve robustness checks include propensity score matching, complexity stratification, and AI/human subgroup analysis.
Results. None of the five hypotheses are supported. After propensity score matching on just-in-time risk profiles, spec'd and unspec'd PRs are indistinguishable in defect introduction (-0.6 percentage points (pp), p = 0.133). Rework shows no protective effect either (+1.2pp, p = 0.001 within-author; +0.5pp, p = 0.146 matched). Specification quality predicts neither defects (p = 0.164) nor rework (p = 0.860), and specifications do not constrain the scope of AI-tagged changes (p = 0.997). The positive raw association between specification artifacts and defects is explained by task complexity: developers spec their hardest work. One subgroup runs the other way - in repositories with no detected AI use, specifications are associated with less rework (-3.3pp, p = 0.014) - but it is one of six subgroup tests and does not survive correction for multiple comparisons.
Conclusion. Specification artifacts proxy for task complexity, not quality improvement. SDD tool adoption itself remains untested: no repository in the sample commits such output.
Cite as
Brenn Hill. Specification Artifacts in Open-Source Pull Request Workflows Do Not Reduce Defects: An Empirical Test of Spec-Driven Development Claims. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 84:1-84:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{hill:LIPIcs.ESEM.2026.84,
author = {Hill, Brenn},
title = {{Specification Artifacts in Open-Source Pull Request Workflows Do Not Reduce Defects: An Empirical Test of Spec-Driven Development Claims}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {84:1--84:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.84},
URN = {urn:nbn:de:0030-drops-280526},
doi = {10.4230/LIPIcs.ESEM.2026.84},
annote = {Keywords: Spec-driven development, specification artifacts, null result, defect prediction, propensity score matching, SZZ, AI-assisted development, within-author fixed effects}
}
Document
Software Engineering in Practice Track Paper
Authors:
Sivajeet Chand, Gedeon Lenz, Patrick Funke, Derui Zhu, and Alexander Pretschner
Abstract
Software migrations are frequent and high-stakes industrial activities. Although large language models (LLMs) are increasingly used in software development, their suitability for migration remains uncertain because migration requires semantic preservation, project-specific context, security, and production accountability. We investigate why practitioners undertake software migrations, what motivates and constrains the use of LLMs for migration, where LLM assistance is considered acceptable, and which factors determine trust in LLM-migrated code. We conducted a practitioner survey with 61 retained industry respondents, combining an internal industrial sample with an open professional survey. We analyzed closed and Likert-style items quantitatively and interpreted open-ended responses thematically. Migration was primarily framed as a bundled quality-risk activity: maintainability, scalability, and security were the dominant motives, while cost reduction rarely appeared as a standalone driver. Respondents valued LLMs mainly as accelerators for exploration, prototyping, and routine transformation, but not as autonomous migration agents. Acceptability was strongly bound by scope and control: 62.3% considered LLMs acceptable most of the time or almost always for single-file migration, compared with only 4.9% for entire-project migration. Trust depended chiefly on semantic correctness, project-context awareness, security, verification evidence, and provenance disclosure. Practitioners do not reject LLM-assisted migration outright; rather, they adopt a conditional trust model. LLMs are viewed as useful migration accelerators, but acceptable production use depends on bounded scope, human review, assurance evidence, traceability, and clear accountability.
Cite as
Sivajeet Chand, Gedeon Lenz, Patrick Funke, Derui Zhu, and Alexander Pretschner. Why Industry Migrates and When It Trusts LLMs: A Practitioner Survey on LLM-Assisted Software Migration. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 85:1-85:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{chand_et_al:LIPIcs.ESEM.2026.85,
author = {Chand, Sivajeet and Lenz, Gedeon and Funke, Patrick and Zhu, Derui and Pretschner, Alexander},
title = {{Why Industry Migrates and When It Trusts LLMs: A Practitioner Survey on LLM-Assisted Software Migration}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {85:1--85:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.85},
URN = {urn:nbn:de:0030-drops-280531},
doi = {10.4230/LIPIcs.ESEM.2026.85},
annote = {Keywords: LLMs, Code Migration, Industrial Insights, LLM-Assisted, Code Translation}
}
Document
Software Engineering in Practice Track Paper
Authors:
Weiyan Sun, Audris Mockus, Brian Ellis, Jun Ge, Madigan Kim, Sahil Kumar, Gursharan Singh, Matt Steiner, Siri Uppalapati, and Nachiappan Nagappan
Abstract
Configuration changes - modifications to feature flags and service parameters - are important to operating services, but they can cause severe outages (SEVs) with substantial service disruption. While code diff risk prediction is well studied, configuration diff risk has received little attention. The code-oriented Diff Risk Score (DRS) model provides limited discriminatory power for config diffs because config risk stems not from code complexity but from the importance, blast radius, and operational readiness of the affected services. We describe the development of CORA (COnfig Risk Analyzer), a dedicated risk model for config diffs. A key insight is that the config-path-to-service mapping bridges from what changed to which services are affected, enabling risk assessment in terms of service criticality rather than code properties.
CORA evolved through three iterations. A logistic regression model with config-specific features achieved an 11% improvement in recall at 5% gating over the DRS model on config diffs. Used for a freeze period, it reduced config gating while outage impact decreased and config diff landing volume increased 91.4%. A unified LightGBM model trained on Meta-wide data achieved a 25.95% improvement in recall at 10% gating, with organization-level improvements ranging from 15% to 66%. An enhanced model with enriched service features and SEV-severity-aware training (having also explored Bayesian hyperparameter optimization) achieved a 22% improvement in recall at 15% gating, intercepting meaningful additional high-severity outage impact in backtesting. To our knowledge, CORA is the first system to apply predictive risk modeling specifically to configuration changes.
Cite as
Weiyan Sun, Audris Mockus, Brian Ellis, Jun Ge, Madigan Kim, Sahil Kumar, Gursharan Singh, Matt Steiner, Siri Uppalapati, and Nachiappan Nagappan. CORA: Config Risk Analyzer - Predicting Risk of Configuration Changes at Scale. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 86:1-86:22, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{sun_et_al:LIPIcs.ESEM.2026.86,
author = {Sun, Weiyan and Mockus, Audris and Ellis, Brian and Ge, Jun and Kim, Madigan and Kumar, Sahil and Singh, Gursharan and Steiner, Matt and Uppalapati, Siri and Nagappan, Nachiappan},
title = {{CORA: Config Risk Analyzer - Predicting Risk of Configuration Changes at Scale}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {86:1--86:22},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.86},
URN = {urn:nbn:de:0030-drops-280542},
doi = {10.4230/LIPIcs.ESEM.2026.86},
annote = {Keywords: Configuration changes, risk prediction, defect prediction, code freeze, software supply chain, gradient boosting, service reliability}
}
Document
Software Engineering in Practice Track Paper
Authors:
Yutan Huang, Jingfan Chen, Fanyu Wang, Chetan Arora, John Grundy, and Cheng Zhang
Abstract
Healthcare AI software teams face growing pressure to translate ethical principles and regulatory obligations into concrete requirements, design decisions, and assurance evidence. Day-to-day requirements engineering (RE) practice still lacks lightweight tool support for connecting abstract frameworks, such as the EU AI Act and the NIST AI Risk Management Framework, to actionable requirements work. This paper builds on a two-phase research project. In the first phase, we conducted an empirical interview study with healthcare AI practitioners spanning clinicians as end users of HealthAI systems and software engineers as developers of HealthAI software systems, which surfaced recurring concerns around transparency, accountability, and the difficulty of mapping regulatory obligations to concrete engineering tasks. Informed by those findings, in the second phase we designed, implemented, and conducted an initial evaluation of the HealthAI Ethics Assistant, an AI-enabled RE tool for healthcare AI software systems.
The tool supports practitioners in generating, validating, and comparing ethical requirements through a structured Create-Validate-Compare workflow. It was implemented as a full-stack web application and grounded in a structured knowledge base combining regulatory guidance (EU AI Act, NIST AI RMF) with practitioner concerns surfaced by the interview study. The work was conducted in close collaboration with an R&D engineer at a medical device manufacturer, who contributed to industrial problem framing and participated as the first practitioner evaluator in the case study reported here.
Our initial evaluation is exploratory, involving a single case study, and used the Technology Acceptance Model and the System Usability Scale. Traceable regulatory references were rated as the most valuable and trustworthy feature, highlighting the importance of explainability and evidence support in compliance-oriented requirements work. The main adoption challenge was not basic usability, but fitting the tool into existing engineering and regulatory workflows. We derive practice-oriented lessons for designing AI-enabled requirements tools in regulated domains: ground LLM outputs in explicit compliance sources, support multiple practitioner perspectives, make generated requirements reviewable rather than authoritative, and align tool use with existing assurance processes. These early results suggest that AI-enabled assistants can help bridge ethical AI principles and practical requirements engineering when designed as auditable, human-in-the-loop support tools rather than autonomous compliance solutions.
Cite as
Yutan Huang, Jingfan Chen, Fanyu Wang, Chetan Arora, John Grundy, and Cheng Zhang. AI Ethics to Requirements Practice: Building and Evaluating the HealthAI Ethics Assistant. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 87:1-87:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{huang_et_al:LIPIcs.ESEM.2026.87,
author = {Huang, Yutan and Chen, Jingfan and Wang, Fanyu and Arora, Chetan and Grundy, John and Zhang, Cheng},
title = {{AI Ethics to Requirements Practice: Building and Evaluating the HealthAI Ethics Assistant}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {87:1--87:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.87},
URN = {urn:nbn:de:0030-drops-280555},
doi = {10.4230/LIPIcs.ESEM.2026.87},
annote = {Keywords: Requirements Engineering, Software Engineering, AI ethics, Healthcare, Large Language Models, Regulatory Compliance.}
}
Document
Software Engineering in Practice Track Paper
Authors:
Fabio Moretti, Simone Ronzoni, Patrizia Scandurra, and Vincenzo Scotti
Abstract
End-to-End (E2E) testing of modern web applications simulates real user interactions to verify the complete application flow, ensuring seamless integration of UI, functionality, and data. It is, in general, costly to create and maintain, especially when the backend is a smart city platform with a data lake and complex user flows that introduce additional complexity. Recent advances in Large Language Models (LLMs) offer new opportunities for automating parts of this process, but their practical adoption in real development environments remains challenging, particularly when privacy, monetary cost, deployment, and control requirements call for local rather than cloud-based models.
This paper investigates the feasibility of introducing local LLMs into the E2E testing process of a real-world smart city web application, the ENEA PELL-IP portal for monitoring public lighting infrastructures. We developed and evaluated GenE2E, a two-stage pipeline that generates test cases from use case specifications and transforms them into executable Playwright tests. The approach was integrated with an existing testing environment and evaluated against manually written baseline tests.
Our findings show that local LLMs can effectively support E2E test generation, but not yet as fully autonomous tools. In our setting, Llama 3.3 outperformed CodeLlama mainly due to its larger context window; single-test generation improved focus and coverage, whereas batch processing reduced model invocations and showed lower generation time in our setup, but the timing results are hardware-dependent and based on a single run; broader efficiency claims require repeated measurements on representative hardware. We also observed that specification quality, completeness of Page Object Model (application-specific API wrapping HTML pages), and use case dependency management had a major impact on generated test quality. Based on this feasibility study, we derive preliminary lessons for practitioners considering local LLMs for E2E test generation.
Cite as
Fabio Moretti, Simone Ronzoni, Patrizia Scandurra, and Vincenzo Scotti. Local LLMs for End-To-End Testing in Practice: Lessons from a Smart City Web Application. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 88:1-88:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{moretti_et_al:LIPIcs.ESEM.2026.88,
author = {Moretti, Fabio and Ronzoni, Simone and Scandurra, Patrizia and Scotti, Vincenzo},
title = {{Local LLMs for End-To-End Testing in Practice: Lessons from a Smart City Web Application}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {88:1--88:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.88},
URN = {urn:nbn:de:0030-drops-280565},
doi = {10.4230/LIPIcs.ESEM.2026.88},
annote = {Keywords: End-to-End testing, Use Case Scenarios, Web applications, Large Language Models}
}
Document
Software Engineering in Practice Track Paper
Authors:
Lo Gullstrand Heander, Agnia Sergeyuk, Ilya Zakharov, Emma Söderberg, and Nikita Mukhortov
Abstract
Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.
Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.
Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.
Results. Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50-3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.
Conclusions. Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention.
Cite as
Lo Gullstrand Heander, Agnia Sergeyuk, Ilya Zakharov, Emma Söderberg, and Nikita Mukhortov. Trust-Calibrated Code Review: A Participatory Design Study of Review Workflows for LLM-Generated Multi-File Changes. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 89:1-89:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{gullstrandheander_et_al:LIPIcs.ESEM.2026.89,
author = {Gullstrand Heander, Lo and Sergeyuk, Agnia and Zakharov, Ilya and S\"{o}derberg, Emma and Mukhortov, Nikita},
title = {{Trust-Calibrated Code Review: A Participatory Design Study of Review Workflows for LLM-Generated Multi-File Changes}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {89:1--89:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.89},
URN = {urn:nbn:de:0030-drops-280578},
doi = {10.4230/LIPIcs.ESEM.2026.89},
annote = {Keywords: code review, participatory design, LLM-generated code, trust calibration, software development tools}
}
Document
Software Engineering in Practice Track Paper
Authors:
Markus Stolze and Mirco Strässle
Abstract
AI-assisted development tools enable software engineers to generate implementations at substantially higher speed and volume than in traditional workflows. Software teams have long relied on guardrails - standing control mechanisms such as code review, linting, testing, and CI/CD pipelines - to maintain quality and coordination. High-throughput AI-assisted generation increases pressure on these guardrails - straining their capacity to keep pace with the volume and rate of generated changes - and reshapes how organizations supervise development workflows, yet relatively little is known about how existing guardrails evolve in response.
We conducted a qualitative interview study with five software engineering practitioners, situated within a broader practitioner survey. Our findings indicate that organizations distribute the work of supervision across multiple guardrail layers: preventive guardrails (produced by externalizing architectural intent and conventions into machine-interpretable form), executable guardrails (linting, testing, and CI/CD repurposed as scalable supervision infrastructure), and human oversight (shifting from line-by-line inspection toward supervisory interpretation focused on architectural reasoning, explainability, and long-term maintainability). We characterize this as a transition from review-centric guardrails toward layered supervision, in which no single guardrail carries the supervision load alone.
Cite as
Markus Stolze and Mirco Strässle. When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 90:1-90:12, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{stolze_et_al:LIPIcs.ESEM.2026.90,
author = {Stolze, Markus and Str\"{a}ssle, Mirco},
title = {{When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {90:1--90:12},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.90},
URN = {urn:nbn:de:0030-drops-280585},
doi = {10.4230/LIPIcs.ESEM.2026.90},
annote = {Keywords: AI-assisted software development, AI coding tools, validation, code review, software teams, guardrails}
}
Document
Software Engineering in Practice Track Paper
Authors:
Jan-Philipp Steghöfer
Abstract
The impact of applying generative AI tools to requirements engineering (RE) in industrial practice remains poorly understood. This paper examines how AI-assisted RE tools are used in industrial practice at XITASO, a medium-sized enterprise for high-tech software engineering, and how they reshape workflows, tool integration, and PO-developer relationships. We combine a 2024 company-wide use-case survey with two rounds of semi-structured interviews with eight product owners (POs) in late 2025 and spring 2026, covering an in-house chatbot and seven commercial AI tools. We identify 15 distinct use cases across four categories: product backlog management, tender management, requirements and domain understanding, and document and artifact creation. Three findings emerge. First, the effect of AI on PO-developer interaction is mixed: the prevailing single-user interaction model can substitute for collaborative dialogue, and developers do not always welcome AI-generated artefacts. Second, tool integration - not tool capability - is the binding constraint: where integration is in place, time savings are dramatic; where it is missing, POs fall back on manual workarounds. Third, AI advances faster than the surrounding organisational systems, so its benefits accrue to individual POs while team processes and customer readiness remain the bottleneck. The empirical GenAI-RE literature remains dominated by early-stage, lab-oriented evaluations of isolated tasks while practice has moved into territory it has not yet studied: practitioners are already assembling cross-tool integrations, navigating customer governance, and renegotiating role boundaries. From these patterns we derive a set of questions practitioners considering AI-assisted RE may ask of their own situation.
Cite as
Jan-Philipp Steghöfer. Faster Than the Team, Faster Than the Customer: Tool Integration, Collaboration, and Organisational Lag in AI-Assisted RE. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 91:1-91:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{steghofer:LIPIcs.ESEM.2026.91,
author = {Stegh\"{o}fer, Jan-Philipp},
title = {{Faster Than the Team, Faster Than the Customer: Tool Integration, Collaboration, and Organisational Lag in AI-Assisted RE}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {91:1--91:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.91},
URN = {urn:nbn:de:0030-drops-280591},
doi = {10.4230/LIPIcs.ESEM.2026.91},
annote = {Keywords: Requirements Engineering, RE in Practice, AI assistance, Backlog Management, Collaboration in RE}
}
Document
Software Engineering in Practice Track Paper
Authors:
Sylwia Kopczyńska, Mirosław Ochodek, Eryk Kosmala, Jędrzej Musiał, Rafał Mroziewski, and Dominik Nitychoruk
Abstract
Requirements Engineering commonly distinguishes between functional and non-functional requirements. In AI-based systems, an orthogonal perspective might also be worth considering: whether requirements concern AI components or not.
This paper reports on a 2.5-year R&D project building MONEKTO, a holistic data-driven recruitment system for a Polish temporary employment agency. The system integrates three data streams: recruitment process events, candidate and job order profiles, and employment histories - to support both candidate-order matching and employee retention prediction.
Drawing on project logs covering June 2023 to January 2026, we identified three categories of requirements in AI-based systems: Non-AI-based requirements, AI-based requirements, and AI-enabling requirements.
We show how these categories differ in their lifecycle: while traditional requirements stabilize after implementation, AI-enabling requirements emerge incrementally and require continuous investment throughout the system lifetime. We report five observations on the emergent nature of AI-enabling requirements and derive seven practitioner lessons for RE in AI-based system projects.
Cite as
Sylwia Kopczyńska, Mirosław Ochodek, Eryk Kosmala, Jędrzej Musiał, Rafał Mroziewski, and Dominik Nitychoruk. AI-Enabling Requirements: An Industrial Experience Report from a Data-Driven Recruitment System. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 92:1-92:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{kopczynska_et_al:LIPIcs.ESEM.2026.92,
author = {Kopczy\'{n}ska, Sylwia and Ochodek, Miros{\l}aw and Kosmala, Eryk and Musia{\l}, J\k{e}drzej and Mroziewski, Rafa{\l} and Nitychoruk, Dominik},
title = {{AI-Enabling Requirements: An Industrial Experience Report from a Data-Driven Recruitment System}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {92:1--92:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.92},
URN = {urn:nbn:de:0030-drops-280604},
doi = {10.4230/LIPIcs.ESEM.2026.92},
annote = {Keywords: requirements engineering, AI-based systems, AI-enabling requirements, industry case study, data-driven recruitment}
}
Document
Software Engineering in Practice Track Paper
Authors:
Audris Mockus, Peter C. Rigby, Don Stewart, Chandra Maddila, and Nachiappan Nagappan
Abstract
AI code-generation tools are widely deployed, yet AI is not applied uniformly: its use varies systematically with the context of the task, the code being modified, and the engineer. Studies that do not account for these contextual differences risk attributing to AI what is actually driven by context. Using data from over 10K developers and 400K diffs at Meta, with character-level provenance tracing of AI-suggested code through editing, review, and landing, we model (a) which context factors predict AI code landing, (b) how landed AI code relates to review time, and (c) its association with production outages (SEVs) - while controlling for contextual confounds. We find that AI code lands more often in less central code, larger changes, test files, and for authors with higher tenure. Review time decreases and SEV rates are lower for diffs with AI code, but both associations are confounded by AI code appearing in less central contexts. These findings demonstrate that naive comparisons of AI vs. non-AI work may reach misleading conclusions, and we provide actionable guidance for controlling these confounds.
Cite as
Audris Mockus, Peter C. Rigby, Don Stewart, Chandra Maddila, and Nachiappan Nagappan. A Preliminary Analysis of the Impact of AI Assisted/Generated Code on Quality, Centrality and Review Time at Scale. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 93:1-93:23, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{mockus_et_al:LIPIcs.ESEM.2026.93,
author = {Mockus, Audris and Rigby, Peter C. and Stewart, Don and Maddila, Chandra and Nagappan, Nachiappan},
title = {{A Preliminary Analysis of the Impact of AI Assisted/Generated Code on Quality, Centrality and Review Time at Scale}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {93:1--93:23},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.93},
URN = {urn:nbn:de:0030-drops-280619},
doi = {10.4230/LIPIcs.ESEM.2026.93},
annote = {Keywords: AI code generation, code review, software quality, context-dependent AI, landing rate}
}
Document
Software Engineering in Practice Track Paper
Authors:
Dennis Schrader, Eva-Maria Schön, Henning Fritzemeier, and Michael Neumann
Abstract
Context. The EU AI Act requires providers and deployers of Artificial Intelligence (AI) systems to implement documentation, risk management, and human oversight. Agile teams that ship AI features in short iterations lack specific artifacts to discharge these duties, since the regulation’s abstract provisions do not map onto agile practices such as the Definition of Done, Sprint Reviews, or working agreements.
Objective. We provide agile teams with an actionable compliance instrument: an evaluated guideline that operationalizes EU AI Act obligations as activities integrable into existing agile practices. We further document the translation method to make its potential adaptation to other regulations examinable.
Method. Following Design Science Research, we assessed each EU AI Act article along three dimensions (technical, normative, organizational). We subsequently classified the articles using a traffic-light scheme and mapped those deemed highly relevant to previously documented pain points of agile teams working with AI. We evaluated the resulting catalog by eliciting practitioners' perceptions through a survey followed by semi-structured interviews with the same participants (N=11) and analyzed the data via qualitative content analysis.
Results. The guideline comprises 12 items covering roles and responsibilities, risk and quality management, transparency and traceability, monitoring, and regulatory sandboxes. Practitioners rated the catalog as understandable and relevant; feasibility varied with organizational maturity. Participants associated perceived adoption feasibility with collective ownership across roles and integration into existing agile events rather than parallel compliance processes.
Conclusions. The catalog gives agile teams a starting point to transform their delivery practices towards EU AI Act compliance without dismantling agile practices. A replication package is publicly available.
Cite as
Dennis Schrader, Eva-Maria Schön, Henning Fritzemeier, and Michael Neumann. Operationalizing the EU AI Act in Agile Software Development: A Guideline-Based Approach. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 94:1-94:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{schrader_et_al:LIPIcs.ESEM.2026.94,
author = {Schrader, Dennis and Sch\"{o}n, Eva-Maria and Fritzemeier, Henning and Neumann, Michael},
title = {{Operationalizing the EU AI Act in Agile Software Development: A Guideline-Based Approach}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {94:1--94:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.94},
URN = {urn:nbn:de:0030-drops-280621},
doi = {10.4230/LIPIcs.ESEM.2026.94},
annote = {Keywords: EU AI Act, AI Governance, Agile Software Development, Regulatory Compliance, Guidelines, Design Science Research}
}
Document
Software Engineering in Practice Track Paper
Authors:
Vekil Bekmyradov, Thomas Bach, Alexander Berndt, Noah C. Puetz, Bartosz Bogacz, and Thomas Bartz-Beielstein
Abstract
Background. Implementing unit tests is an important yet time-consuming activity in software development. Therefore, Large Language Models (LLMs) are being used increasingly often for generating unit tests. Many previous work report promising results for open-source repositories written in common languages such as Python or Java. However, the performance of LLM-based test generation in large proprietary software projects written in C++ remains largely unknown.
Aims. To reduce this gap, we investigate the performance of LLM-based test generation on a large closed-source C++ software project, which is not part of the training corpus of LLMs.
Method. We study test generation on SAP HANA, a large C++ DBMS, and compare it to the open-source key-value store LevelDB. We analyze two LLM-based test generation approaches, each including an iterative repair loop for compilation failures, along five metrics: compilation success rate (CSR), execution success rate, line coverage, branch coverage, and mutation score (MS).
Results. The metrics show a considerable gap between the two systems. The LLM-generated tests for LevelDB match the quality of the existing tests. On SAP HANA, the LLM-generated tests achieve only 25.2% MS and underwhelming coverage results. Iterative repair with at least 3 iterations is important, reaching 97.6% CSR on SAP HANA after 10 iterations. However, with more iterations, LLM may prioritize compilability over test quality, resulting in tests with weak assertions.
Conclusions. Results reported on open-source projects, whose code is often part of LLM training data, may not generalize to unseen, industry-grade C++ codebases. The effectiveness of iterative repair saturates after several iterations. When applying LLMs to source code not in training data, practitioners should carefully select the context provided to the models.
Cite as
Vekil Bekmyradov, Thomas Bach, Alexander Berndt, Noah C. Puetz, Bartosz Bogacz, and Thomas Bartz-Beielstein. Evaluating LLM-Based Test Generation for a Large Industrial C++ Database System. A Case Study on SAP HANA. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 95:1-95:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)
Copy BibTex To Clipboard
@InProceedings{bekmyradov_et_al:LIPIcs.ESEM.2026.95,
author = {Bekmyradov, Vekil and Bach, Thomas and Berndt, Alexander and Puetz, Noah C. and Bogacz, Bartosz and Bartz-Beielstein, Thomas},
title = {{Evaluating LLM-Based Test Generation for a Large Industrial C++ Database System. A Case Study on SAP HANA}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {95:1--95:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.95},
URN = {urn:nbn:de:0030-drops-280632},
doi = {10.4230/LIPIcs.ESEM.2026.95},
annote = {Keywords: LLM, Test Generation, Software Testing, Mutation Testing, DBMS}
}