Search Results

Documents authored by Naqvi, Syed Fakhar Abbas


Artifact
Software
Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks

Authors: Syed Fakhar Abbas Naqvi and Hina Anwar


Abstract

Cite as

Syed Fakhar Abbas Naqvi, Hina Anwar. Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks (Software, Source Code). Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@misc{dagstuhl-artifact-28073,
   title = {{Benchmarking Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks}}, 
   author = {Naqvi, Syed Fakhar Abbas and Anwar, Hina},
   note = {Software, swhId: \href{https://archive.softwareheritage.org/swh:1:dir:e2f5ea38e5e8b633db8ef0baf818d58b194468a9;origin=https://github.com/syedfakhar25/empirical-llm-bugfix-benchmark;visit=swh:1:snp:30dfb758d4b717c0deb55403872688c9af37f075;anchor=swh:1:rev:e3cc68f625c69eacd334437f1c9c1343aae6b348}{\texttt{swh:1:dir:e2f5ea38e5e8b633db8ef0baf818d58b194468a9}} (visited on 2026-10-05)},
   url = {https://github.com/syedfakhar25/empirical-llm-bugfix-benchmark/tree/main},
   doi = {10.4230/artifacts.28073},
}
Document
Technical Track Paper
Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks

Authors: Syed Fakhar Abbas Naqvi and Hina Anwar

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Large Language Models (LLMs) are increasingly used in software engineering tasks such as code generation, bug detection, and program repair. However, their use introduces inference-time costs that are rarely considered together with functional performance, especially in debugging and repair workflows where models may be invoked repeatedly. Aim. This study evaluates the correctness-energy trade-offs of LLMs across model scales in Python bug-fixing tasks. We examine how model size relates to code correctness, inference time, and GPU energy consumption. Method. We conduct a controlled empirical evaluation of six open-source code-oriented LLMs, ranging from 1.5B to 15B parameters, on 40 real-world Python bugs from the BugsInPy benchmark. Each model is executed multiple times per task under a consistent hardware and prompting setup. Repair performance is measured using Pass@k and test-suite outcomes, while computational cost is measured using inference time and GPU energy consumption. We also analyze generated outputs to identify recurring failure patterns. Results. Increasing model size substantially raised computational cost without consistently improving correctness. The smallest model, Qwen 1.5B, achieved the best overall performance, with a Pass@1 score of 43% and an average runtime of 38.7 seconds. In contrast, StarCoder 15B achieved a Pass@1 score of 23%, required 96.7 seconds, and consumed approximately 7 times more energy per inference while solving fewer tasks. Higher success rates were also observed in web-related projects, suggesting that structurally localized bugs are more effectively handled by current models. Conclusion. On the evaluated localized Python bug-fixing tasks, smaller LLMs offered a more favorable trade-off between correctness, inference time, and energy consumption. These results suggest that scaling model size does not automatically lead to more efficient LLM-based repair, and that inference cost should be considered alongside functional correctness when comparing models.

Cite as

Syed Fakhar Abbas Naqvi and Hina Anwar. Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 34:1-34:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{naqvi_et_al:LIPIcs.ESEM.2026.34,
  author =	{Naqvi, Syed Fakhar Abbas and Anwar, Hina},
  title =	{{Bigger Is Not Always Better: Performance and Sustainability Trade-Offs in LLMs for Python Bug-Fixing Tasks}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{34:1--34:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.34},
  URN =		{urn:nbn:de:0030-drops-280023},
  doi =		{10.4230/LIPIcs.ESEM.2026.34},
  annote =	{Keywords: Large Language Models, Bug Fixing, Energy Consumption, Empirical Software Engineering, Sustainability}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail