<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:32:40Z</responseDate>
  <request identifier="27981" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:27981</identifier>
        <datestamp>2026-10-05T06:44:02Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>Characterizing Feedback Statements in Machine Learning Jupyter Notebooks</dc:title>
          <dc:creator>Shome, Arumoy</dc:creator>
          <dc:creator>Cruz, Luís</dc:creator>
          <dc:creator>Spinellis, Diomidis</dc:creator>
          <dc:creator>van Deursen, Arie</dc:creator>
          <dc:subject>Empirical software engineering</dc:subject>
          <dc:subject>machine learning</dc:subject>
          <dc:subject>Jupyter notebooks</dc:subject>
          <dc:subject>software testing</dc:subject>
          <dc:subject>assertions</dc:subject>
          <dc:subject>mining software repositories</dc:subject>
          <dc:description>Background. Machine learning development in Jupyter notebooks is iterative and feedback-driven. Practitioners author statements that reveal information about program execution and use this information to decide what to do next. We call these feedback statements and identify two forms: exploratory statements that display values for visual inspection and validation statements that enforce conditions programmatically through assertions.&#13;
&#13;
Aims. Many failures in ML systems do not surface as exceptions, and consequently escape the crash-based analyses that dominate prior empirical work on ML notebooks. This study examines what practitioners check to catch the failures that would otherwise pass silently, by characterizing feedback statements that encode the practitioner’s mental model of what the code should do and what could go wrong.&#13;
&#13;
Method. We mine 297,851 publicly available Python Jupyter notebooks from GitHub and Kaggle, and extract 1,092,780 feedback statements. We sample 816 statements through proportional stratified sampling from semantic clusters obtained from CodeBERT embeddings, and apply grounded theory and open coding to manually label and analyze each statement.&#13;
&#13;
Results. We contribute a taxonomy of feedback statements in ML Jupyter notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears. The taxonomy reveals that feedback in ML notebooks is overwhelmingly exploratory, and that the two platforms host qualitatively different modes of ML work. We further map our taxonomy to an existing crash taxonomy and find that our taxonomy captures defensive practices against silent failures that crash analysis cannot observe.&#13;
&#13;
Conclusions. Our findings indicate that notebook source should be treated as a confounder in studies of ML developer practice, surface opportunities for notebook tooling, and motivate empirical study of silent ML failures. We release the corpus of 1,092,780 feedback statements and the codebook, to support replication and tooling research.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Arumoy Shome and Luís Cruz and Diomidis Spinellis and Arie van Deursen</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.13</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-279814</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.13</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
