,
Thomas Bach
,
Alexander Berndt
,
Noah C. Puetz
,
Bartosz Bogacz
,
Thomas Bartz-Beielstein
Creative Commons Attribution 4.0 International license
Background. Implementing unit tests is an important yet time-consuming activity in software development. Therefore, Large Language Models (LLMs) are being used increasingly often for generating unit tests. Many previous work report promising results for open-source repositories written in common languages such as Python or Java. However, the performance of LLM-based test generation in large proprietary software projects written in C++ remains largely unknown. Aims. To reduce this gap, we investigate the performance of LLM-based test generation on a large closed-source C++ software project, which is not part of the training corpus of LLMs. Method. We study test generation on SAP HANA, a large C++ DBMS, and compare it to the open-source key-value store LevelDB. We analyze two LLM-based test generation approaches, each including an iterative repair loop for compilation failures, along five metrics: compilation success rate (CSR), execution success rate, line coverage, branch coverage, and mutation score (MS). Results. The metrics show a considerable gap between the two systems. The LLM-generated tests for LevelDB match the quality of the existing tests. On SAP HANA, the LLM-generated tests achieve only 25.2% MS and underwhelming coverage results. Iterative repair with at least 3 iterations is important, reaching 97.6% CSR on SAP HANA after 10 iterations. However, with more iterations, LLM may prioritize compilability over test quality, resulting in tests with weak assertions. Conclusions. Results reported on open-source projects, whose code is often part of LLM training data, may not generalize to unseen, industry-grade C++ codebases. The effectiveness of iterative repair saturates after several iterations. When applying LLMs to source code not in training data, practitioners should carefully select the context provided to the models.
@InProceedings{bekmyradov_et_al:LIPIcs.ESEM.2026.95,
author = {Bekmyradov, Vekil and Bach, Thomas and Berndt, Alexander and Puetz, Noah C. and Bogacz, Bartosz and Bartz-Beielstein, Thomas},
title = {{Evaluating LLM-Based Test Generation for a Large Industrial C++ Database System. A Case Study on SAP HANA}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {95:1--95:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.95},
URN = {urn:nbn:de:0030-drops-280632},
doi = {10.4230/LIPIcs.ESEM.2026.95},
annote = {Keywords: LLM, Test Generation, Software Testing, Mutation Testing, DBMS}
}