Search Results

Documents authored by Anwar, Hina


Artifact
Software
Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks

Authors: Syed Fakhar Abbas Naqvi and Hina Anwar


Abstract

Cite as

Syed Fakhar Abbas Naqvi, Hina Anwar. Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks (Software, Source Code). Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@misc{dagstuhl-artifact-28073,
   title = {{Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks}}, 
   author = {Naqvi, Syed Fakhar Abbas and Anwar, Hina},
   note = {Software, swhId: \href{https://archive.softwareheritage.org/swh:1:dir:e2f5ea38e5e8b633db8ef0baf818d58b194468a9;origin=https://github.com/syedfakhar25/empirical-llm-bugfix-benchmark;visit=swh:1:snp:30dfb758d4b717c0deb55403872688c9af37f075;anchor=swh:1:rev:e3cc68f625c69eacd334437f1c9c1343aae6b348}{\texttt{swh:1:dir:e2f5ea38e5e8b633db8ef0baf818d58b194468a9}} (visited on 2026-10-05)},
   url = {https://github.com/syedfakhar25/empirical-llm-bugfix-benchmark/tree/main},
   doi = {10.4230/artifacts.28073},
}
Document
Technical Track Paper
Balancing Green and Clean Code: Prompting for LLM-Based Refactoring

Authors: Feray Gulu-zada and Hina Anwar

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. LLMs are increasingly used for code refactoring, but their effects on multiple software quality dimensions remain unclear. Green coding targets energy efficiency, while clean coding emphasizes readability and maintainability. These objectives may interact or conflict, making LLM-based refactoring a multi-objective problem. Aims. This study investigates whether prompting can guide LLM-based refactoring toward balanced green and clean code outcomes. We examine how different prompting strategies affect energy consumption, runtime performance, and maintainability. Method. We conduct a controlled empirical study on 15 human-written Python scripts using four LLMs and five prompting strategies: clean, green, multi-objective, clean-to-green staged, and green-to-clean staged prompting. Generated refactorings are validated for correctness before analysis. Valid outputs are evaluated using CPU energy consumption, runtime, maintainability index, cyclomatic complexity, source lines of code, and Halstead volume. We analyze results using descriptive statistics, statistical tests, correlation analysis, bootstrap confidence intervals, and Pareto-efficiency analysis. Results. Of 300 generated refactorings, 267 passed validation. Among valid outputs, 53.18% (142 of 267) reduced energy consumption, but the overall energy effect was small, variable, and not statistically significant across prompting strategies. Energy and runtime changes showed a moderate positive correlation, while maintainability showed weaker relationships with efficiency metrics. Prompting strategies significantly affected maintainability but not energy or runtime. Clean and multi-objective prompting produced the most balanced Pareto outcomes, though these differences were mainly driven by maintainability. Conclusions. LLM-based refactoring can sometimes reduce energy consumption without necessarily degrading maintainability, but its effects are inconsistent and depend on model choice, prompt design, and correctness validation. Prompting is useful for steering LLM-based refactoring, but current evidence supports cautious, validated use rather than assuming reliable multi-objective optimization.

Cite as

Feray Gulu-zada and Hina Anwar. Balancing Green and Clean Code: Prompting for LLM-Based Refactoring. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 29:1-29:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{guluzada_et_al:LIPIcs.ESEM.2026.29,
  author =	{Gulu-zada, Feray and Anwar, Hina},
  title =	{{Balancing Green and Clean Code: Prompting for LLM-Based Refactoring}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{29:1--29:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.29},
  URN =		{urn:nbn:de:0030-drops-279971},
  doi =		{10.4230/LIPIcs.ESEM.2026.29},
  annote =	{Keywords: Large language models, code refactoring, energy efficiency, maintainability, green software, prompt engineering}
}
Document
Technical Track Paper
Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks

Authors: Syed Fakhar Abbas Naqvi and Hina Anwar

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Large Language Models (LLMs) are increasingly used in software engineering tasks such as code generation, bug detection, and program repair. However, their use introduces inference-time costs that are rarely considered together with functional performance, especially in debugging and repair workflows where models may be invoked repeatedly. Aim. This study evaluates the correctness-energy trade-offs of LLMs across model scales in Python bug-fixing tasks. We examine how model size relates to code correctness, inference time, and GPU energy consumption. Method. We conduct a controlled empirical evaluation of six open-source code-oriented LLMs, ranging from 1.5B to 15B parameters, on 40 real-world Python bugs from the BugsInPy benchmark. Each model is executed multiple times per task under a consistent hardware and prompting setup. Repair performance is measured using Pass@k and test-suite outcomes, while computational cost is measured using inference time and GPU energy consumption. We also analyze generated outputs to identify recurring failure patterns. Results. Increasing model size substantially raised computational cost without consistently improving correctness. The smallest model, Qwen 1.5B, achieved the best overall performance, with a Pass@1 score of 43% and an average runtime of 38.7 seconds. In contrast, StarCoder 15B achieved a Pass@1 score of 23%, required 96.7 seconds, and consumed approximately 7 times more energy per inference while solving fewer tasks. Higher success rates were also observed in web-related projects, suggesting that structurally localized bugs are more effectively handled by current models. Conclusion. On the evaluated localized Python bug-fixing tasks, smaller LLMs offered a more favorable trade-off between correctness, inference time, and energy consumption. These results suggest that scaling model size does not automatically lead to more efficient LLM-based repair, and that inference cost should be considered alongside functional correctness when comparing models.

Cite as

Syed Fakhar Abbas Naqvi and Hina Anwar. Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 34:1-34:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{naqvi_et_al:LIPIcs.ESEM.2026.34,
  author =	{Naqvi, Syed Fakhar Abbas and Anwar, Hina},
  title =	{{Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{34:1--34:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.34},
  URN =		{urn:nbn:de:0030-drops-280023},
  doi =		{10.4230/LIPIcs.ESEM.2026.34},
  annote =	{Keywords: Large Language Models, Bug Fixing, Energy Consumption, Empirical Software Engineering, Sustainability}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail