,
Regina Hebig
Creative Commons Attribution 4.0 International license
Background. Generating software tests and implementations with Large Language Models (LLMs) is becoming more common. However, using the same LLM for the generation of implementation and tests might make both suffer from the same biases, potentially decreasing the tests' ability to detect faults compared to tests generated by a different LLM. Aims. In this paper, we want to investigate this notion by combining different LLMs for generating implementation and tests. Method. We generate Elixir code and test cases with 5 LLMs and compare them to a baseline of manually written code and tests. We wrote and generated 459 distinct implementations and 2709 test suites, manually analyzing 9417 test case failures. Results. We find no difference between tests generated by the same or a different LLM as the code. However, of the 247 implementations flagged as faulty, 99 were only flagged by manually written tests. Conclusions. Our results indicate that LLM-generated tests alone are not yet sufficient.
@InProceedings{leu_et_al:LIPIcs.ESEM.2026.11,
author = {Leu, Noah and Oertel, Julian and Hebig, Regina},
title = {{Don't Bother to Use a Second LLM and Write Tests Yourself! A Study on Elixir}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {11:1--11:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.11},
URN = {urn:nbn:de:0030-drops-279793},
doi = {10.4230/LIPIcs.ESEM.2026.11},
annote = {Keywords: LLMs, Code Generation, Test Generation, Rare Languages}
}