,
Tatsuya Shirai
,
Yutaro Kashiwa
,
Masanari Kondo
,
Yasutaka Kamei
,
Hajimu Iida
Creative Commons Attribution 4.0 International license
Background. Continuous Integration (CI) aims to shorten release cycles by automating tests on every change. When the same test method keeps failing across consecutive revisions, new defects get buried among existing failures, dulling developer vigilance and weakening CI’s benefit of rapid bug localization. Aims. We define Test Alert Snooze as the state in which the same test method fails across two or more consecutive revisions, and provide its first empirical characterization: its prevalence, persistence in revisions and elapsed time, the commits developers make during Test Alert Snooze, and what they change to resolve it. Method. We analyzed 27 open-source Python and Java projects on GitHub Actions with high test execution frequencies. From build histories and logs over a 90-day window (December 21, 2025 to March 20, 2026), we identified test methods that failed across two or more consecutive revisions, and classified both the commits made during Test Alert Snooze and those that resolved it. Results. Test Alert Snooze accounts for 42.9% of observed test failures, with a median persistence of 2 consecutive revisions and about one day; extreme cases reached 11 commits and 16 days. During Test Alert Snooze, fix commits made up only 8.2% of developer activity, while docs, test, refactor, and feat collectively dominated. Resolutions were mostly single-category modifications to Test or Product files, and among resolving commits test was the most frequent (32.4%), well above fix (11.8%). Conclusions. Test-failure persistence is not a single phenomenon but differs across the build, job, and test-method levels. At the test-method level, resolving a failure often takes the form of test-code maintenance rather than bug fixing, so predicting, detecting, and repairing Test Alert Snooze calls for techniques targeting test code alongside product code. This first characterization lays the groundwork for such techniques and for CI features that visualize test-method-level failure persistence.
@InProceedings{shirakawa_et_al:LIPIcs.ESEM.2026.24,
author = {Shirakawa, Ayane and Shirai, Tatsuya and Kashiwa, Yutaro and Kondo, Masanari and Kamei, Yasutaka and Iida, Hajimu},
title = {{Test Alert Snooze: An Empirical Study of Consecutive Test Failures on CI}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {24:1--24:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.24},
URN = {urn:nbn:de:0030-drops-279929},
doi = {10.4230/LIPIcs.ESEM.2026.24},
annote = {Keywords: Continuous integration (CI), Test failures, Empirical analysis}
}