,
Hina Anwar
Creative Commons Attribution 4.0 International license
Background. LLMs are increasingly used for code refactoring, but their effects on multiple software quality dimensions remain unclear. Green coding targets energy efficiency, while clean coding emphasizes readability and maintainability. These objectives may interact or conflict, making LLM-based refactoring a multi-objective problem. Aims. This study investigates whether prompting can guide LLM-based refactoring toward balanced green and clean code outcomes. We examine how different prompting strategies affect energy consumption, runtime performance, and maintainability. Method. We conduct a controlled empirical study on 15 human-written Python scripts using four LLMs and five prompting strategies: clean, green, multi-objective, clean-to-green staged, and green-to-clean staged prompting. Generated refactorings are validated for correctness before analysis. Valid outputs are evaluated using CPU energy consumption, runtime, maintainability index, cyclomatic complexity, source lines of code, and Halstead volume. We analyze results using descriptive statistics, statistical tests, correlation analysis, bootstrap confidence intervals, and Pareto-efficiency analysis. Results. Of 300 generated refactorings, 267 passed validation. Among valid outputs, 53.18% (142 of 267) reduced energy consumption, but the overall energy effect was small, variable, and not statistically significant across prompting strategies. Energy and runtime changes showed a moderate positive correlation, while maintainability showed weaker relationships with efficiency metrics. Prompting strategies significantly affected maintainability but not energy or runtime. Clean and multi-objective prompting produced the most balanced Pareto outcomes, though these differences were mainly driven by maintainability. Conclusions. LLM-based refactoring can sometimes reduce energy consumption without necessarily degrading maintainability, but its effects are inconsistent and depend on model choice, prompt design, and correctness validation. Prompting is useful for steering LLM-based refactoring, but current evidence supports cautious, validated use rather than assuming reliable multi-objective optimization.
@InProceedings{guluzada_et_al:LIPIcs.ESEM.2026.29,
author = {Gulu-zada, Feray and Anwar, Hina},
title = {{Balancing Green and Clean Code: Prompting for LLM-Based Refactoring}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {29:1--29:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.29},
URN = {urn:nbn:de:0030-drops-279971},
doi = {10.4230/LIPIcs.ESEM.2026.29},
annote = {Keywords: Large language models, code refactoring, energy efficiency, maintainability, green software, prompt engineering}
}