Search Results

Documents authored by Cruz, Luís


Document
Technical Track Paper
Characterizing Feedback Statements in Machine Learning Jupyter Notebooks

Authors: Arumoy Shome, Luís Cruz, Diomidis Spinellis, and Arie van Deursen

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Machine learning development in Jupyter notebooks is iterative and feedback-driven. Practitioners author statements that reveal information about program execution and use this information to decide what to do next. We call these feedback statements and identify two forms: exploratory statements that display values for visual inspection and validation statements that enforce conditions programmatically through assertions. Aims. Many failures in ML systems do not surface as exceptions, and consequently escape the crash-based analyses that dominate prior empirical work on ML notebooks. This study examines what practitioners check to catch the failures that would otherwise pass silently, by characterizing feedback statements that encode the practitioner’s mental model of what the code should do and what could go wrong. Method. We mine 297,851 publicly available Python Jupyter notebooks from GitHub and Kaggle, and extract 1,092,780 feedback statements. We sample 816 statements through proportional stratified sampling from semantic clusters obtained from CodeBERT embeddings, and apply grounded theory and open coding to manually label and analyze each statement. Results. We contribute a taxonomy of feedback statements in ML Jupyter notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears. The taxonomy reveals that feedback in ML notebooks is overwhelmingly exploratory, and that the two platforms host qualitatively different modes of ML work. We further map our taxonomy to an existing crash taxonomy and find that our taxonomy captures defensive practices against silent failures that crash analysis cannot observe. Conclusions. Our findings indicate that notebook source should be treated as a confounder in studies of ML developer practice, surface opportunities for notebook tooling, and motivate empirical study of silent ML failures. We release the corpus of 1,092,780 feedback statements and the codebook, to support replication and tooling research.

Cite as

Arumoy Shome, Luís Cruz, Diomidis Spinellis, and Arie van Deursen. Characterizing Feedback Statements in Machine Learning Jupyter Notebooks. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 13:1-13:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{shome_et_al:LIPIcs.ESEM.2026.13,
  author =	{Shome, Arumoy and Cruz, Lu{\'\i}s and Spinellis, Diomidis and van Deursen, Arie},
  title =	{{Characterizing Feedback Statements in Machine Learning Jupyter Notebooks}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{13:1--13:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.13},
  URN =		{urn:nbn:de:0030-drops-279814},
  doi =		{10.4230/LIPIcs.ESEM.2026.13},
  annote =	{Keywords: Empirical software engineering, machine learning, Jupyter notebooks, software testing, assertions, mining software repositories}
}
Document
Technical Track Paper
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

Authors: Enrique Barba Roque, Luís Cruz, and Annibale Panichella

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their high computational demands and energy consumption raise sustainability concerns and hinder their use on consumer hardware and resource-constrained platforms. A common way to report the computational cost of an LLM in the literature and industry is to use the number of Floating Point Operations (FLOPs) required to perform a pass over the network. Aims. This paper investigates the implications of energy-aware knowledge distillation for SE, aiming to improve model efficiency while maintaining performance and to determine whether FLOPs is a reliable energy-aware metric. Method. We conduct a controlled experiment using Morph, a Many-Objective Optimization-based distillation methodology, to empirically examine whether FLOPs accurately reflect energy consumption in Clone Detection and Vulnerability Prediction tasks. We extend this methodology to include energy-surrogate models that directly estimate CPU and GPU energy consumption during optimization, and we apply Morph to generative tasks using CodeT5+ for code summarization. Results. Our results show that FLOPs is not always a reliable indicator of energy consumption, and better results can be achieved by using energy-surrogate models. Distilled student models can reduce inference energy consumption by up to 90% and memory usage by 86%, with only modest accuracy trade-offs. Conclusions. Energy-aware knowledge distillation when guided by direct energy surrogates rather than FLOPs can improve the energy consumption, sustainability, and deployability of LLMs for SE applications, enabling efficient models on consumer hardware.

Cite as

Enrique Barba Roque, Luís Cruz, and Annibale Panichella. Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 22:1-22:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{barbaroque_et_al:LIPIcs.ESEM.2026.22,
  author =	{Barba Roque, Enrique and Cruz, Lu{\'\i}s and Panichella, Annibale},
  title =	{{Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{22:1--22:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.22},
  URN =		{urn:nbn:de:0030-drops-279905},
  doi =		{10.4230/LIPIcs.ESEM.2026.22},
  annote =	{Keywords: Knowledge distillation, Green AI, Many-objective Optimization, LLMs for Code, FLOPs, AI for SE}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail