,
Faezeh Amou Najafabadi
,
Ilias Gerostathopoulos
,
Patricia Lago
Creative Commons Attribution 4.0 International license
Background. Automated Machine Learning (AutoML) is increasingly embedded in software and data engineering workflows, automating the design and optimization of machine learning (ML) pipelines. AutoML has been extensively synthesized through secondary studies that map tools, algorithms, and evaluation practices. However, these reviews predominantly reflect a research-centric perspective and provide limited evidence on how AutoML is actually used, adapted, and experienced in practice. Traditional surveys and interviews capture only a narrow slice of the practitioner community, leaving open how well academic claims about AutoML align with day-to-day engineering realities. Aims. This study aims to (i) assess to what extent AutoML practices and challenges reported in existing secondary studies reflect practitioners' experiences, and (ii) identify mismatches between research and practice by empirically comparing academic claims with practitioner experiences shared on YouTube. Method. Following established guidelines for multivocal literature reviews in software engineering, we conduct a multivocal evidence synthesis combining 16 peer-reviewed secondary studies (systematic and multivocal literature reviews) with 30 practitioner-created YouTube videos in which engineers, data scientists, and ML practitioners describe their AutoML workflows, workarounds, and pain points. Results. Our analysis reveals a marked imbalance: academic reviews emphasize algorithm selection, search strategies, and benchmarking, whereas practitioners foreground data quality and cleaning, resource and cost constraints, tool robustness, scalability, and integration into production pipelines. We also observe areas of alignment, for example around hyperparameter optimization and the need for reproducible workflows. Conclusions. Treating YouTube videos as curated gray literature under established multivocal review guidelines, our study shows that they provide complementary, practice-grounded evidence about AutoML adoption. The identified gaps and alignments suggest concrete directions for empirical AutoML research that more directly address operational and organizational constraints in real-world software engineering settings.
@InProceedings{rajenthiram_et_al:LIPIcs.ESEM.2026.27,
author = {Rajenthiram, Keerthiga and Najafabadi, Faezeh Amou and Gerostathopoulos, Ilias and Lago, Patricia},
title = {{A Multivocal Review of AutoML Practices, Challenges, Opportunities and Open Issues}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {27:1--27:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.27},
URN = {urn:nbn:de:0030-drops-279953},
doi = {10.4230/LIPIcs.ESEM.2026.27},
annote = {Keywords: Automated Machine Learning (AutoML), Multivocal Literature Review, YouTube Videos, Gray Literature, Empirical Study, Research–Practice Gap}
}