Search Results

Documents authored by Iida, Hajimu


Document
Technical Track Paper
Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI

Authors: Ayane Shirakawa, Tatsuya Shirai, Yutaro Kashiwa, Masanari Kondo, Yasutaka Kamei, and Hajimu Iida

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Continuous Integration (CI) aims to shorten release cycles by automating tests on every change. When the same test method keeps failing across consecutive revisions, new defects get buried among existing failures, dulling developer vigilance and weakening CI’s benefit of rapid bug localization. Aims. We define Test Alert Snooze as the state in which the same test method fails across two or more consecutive revisions, and provide its first empirical characterization: its prevalence, persistence in revisions and elapsed time, the commits developers make during Test Alert Snooze, and what they change to resolve it. Method. We analyzed 27 open-source Python and Java projects on GitHub Actions with high test execution frequencies. From build histories and logs over a 90-day window (December 21, 2025 to March 20, 2026), we identified test methods that failed across two or more consecutive revisions, and classified both the commits made during Test Alert Snooze and those that resolved it. Results. Test Alert Snooze accounts for 42.9% of observed test failures, with a median persistence of 2 consecutive revisions and about one day; extreme cases reached 11 commits and 16 days. During Test Alert Snooze, fix commits made up only 8.2% of developer activity, while docs, test, refactor, and feat collectively dominated. Resolutions were mostly single-category modifications to Test or Product files, and among resolving commits test was the most frequent (32.4%), well above fix (11.8%). Conclusions. Test-failure persistence is not a single phenomenon but differs across the build, job, and test-method levels. At the test-method level, resolving a failure often takes the form of test-code maintenance rather than bug fixing, so predicting, detecting, and repairing Test Alert Snooze calls for techniques targeting test code alongside product code. This first characterization lays the groundwork for such techniques and for CI features that visualize test-method-level failure persistence.

Cite as

Ayane Shirakawa, Tatsuya Shirai, Yutaro Kashiwa, Masanari Kondo, Yasutaka Kamei, and Hajimu Iida. Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 24:1-24:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{shirakawa_et_al:LIPIcs.ESEM.2026.24,
  author =	{Shirakawa, Ayane and Shirai, Tatsuya and Kashiwa, Yutaro and Kondo, Masanari and Kamei, Yasutaka and Iida, Hajimu},
  title =	{{Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{24:1--24:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.24},
  URN =		{urn:nbn:de:0030-drops-279929},
  doi =		{10.4230/LIPIcs.ESEM.2026.24},
  annote =	{Keywords: Continuous integration (CI), Test failures, Empirical analysis}
}
Document
Technical Track Paper
Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-Change, and Impact

Authors: Suzuka Yoshimoto, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, and Hajimu Iida

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. When developers write a TODO or FIXME comment, they are explicitly admitting that the code is suboptimal: a built-in warning that this logic deserves extra scrutiny. Yet it is an open question whether Self-Admitted Technical Debt (SATD) actually receives that scrutiny in the form of software testing. Aim. We aim to characterize the relationship between SATD and testing across three dimensions: the extent to which SATD-affected code is covered by existing tests, whether developers synchronize test additions with debt resolution, and whether such testing affects the long-term observability of resulting defects. Method. For that, we conducted an empirical study on eight open-source Java projects, analyzing test coverage of 784 SATD instances identified in the latest releases and performing a longitudinal examination of 5,175 SATD removal events. Results. Our results show that while 60.7% of SATD-affected code is covered by existing test suites, developers rarely synchronize test modifications with debt resolution; manual inspection confirms that only 3.4% of SATD removal commits include new tests specifically targeting the resolved debt (vs. 12.5% that co-add tests in the same commit). Longitudinal analysis further suggests that SATD resolutions exhibit nearly identical localized bug induction rates within short-to-medium-term windows regardless of test modifications. However, over a longer, unrestricted observation window, a slight divergence emerges where the test-added group reaches a higher cumulative defect alignment probability (6.32% vs. 4.37%), a counterintuitive trend potentially driven by the selective testing of inherently complex components. Conclusion. Developers treat SATD repayment as an ordinary code change rather than as a high-risk maintenance activity: most debt removals proceed without targeted verification, despite the developer’s own prior flag that the code is suboptimal.

Cite as

Suzuka Yoshimoto, Kosei Horikawa, Daniel Feitosa, Yutaro Kashiwa, and Hajimu Iida. Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-Change, and Impact. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 36:1-36:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{yoshimoto_et_al:LIPIcs.ESEM.2026.36,
  author =	{Yoshimoto, Suzuka and Horikawa, Kosei and Feitosa, Daniel and Kashiwa, Yutaro and Iida, Hajimu},
  title =	{{Is Self-Admitted Technical Debt Tested? An Empirical Study of Coverage, Co-Change, and Impact}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{36:1--36:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.36},
  URN =		{urn:nbn:de:0030-drops-280048},
  doi =		{10.4230/LIPIcs.ESEM.2026.36},
  annote =	{Keywords: Self-Admitted Technical Debt, Test Coverage, Software Testing}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail