,
Yanjun Wu
Creative Commons Attribution 4.0 International license
Background. Targeted unit tests, which aim to reveal a bug at a specific location in the program, are useful for various software development tasks, such as static analysis alarm validation, crash reproduction, or patch testing. While automated unit test generation has progressed significantly, existing approaches remain inefficient for targeted tests due to their reliance on manually-crafted heuristics. To address this shortcoming, Large Language Model (LLM) agents emerge as a potent solution that is free from predefined heuristics, owing to their superior reasoning and tool-use capabilities. Moreover, prior studies have shown that LLM agents can achieve comparable performance across several software engineering tasks with only minor adaptations. However, whether this adaptability extends to targeted unit test generation remains unexplored. Aims. This study aims to investigate the capabilities and limitations of LLM agents in targeted unit test generation under minimal custom setup. Method. We evaluate various combinations of frontier agent scaffolds and models on a benchmark featuring 210 real-world Java runtime exception bugs across 66 software projects. Specifically, LLM agents are tasked with generating a unit test suite that reproduce a target bug given only its location and the corresponding codebase. Results. Our evaluation shows that the top-performing combinations achieve a 89.5% success rate in target bug reproduction, surpassing the state-of-the-art unit test generation tools. Nevertheless, we identify recurring failure modes that illustrate current LLM agents' blind spots. Conclusions. These findings underscore the potential of LLM agents in targeted unit test generation, and provide valuable insights for further research in this area.
@InProceedings{liu_et_al:LIPIcs.ESEM.2026.47,
author = {Liu, Jiayu and Wu, Yanjun},
title = {{Beyond Heuristics? Rethinking Targeted Unit Test Generation in the Era of LLM Agents}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {47:1--47:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.47},
URN = {urn:nbn:de:0030-drops-280156},
doi = {10.4230/LIPIcs.ESEM.2026.47},
annote = {Keywords: Large Language Model, Agents, Unit Test Generation, Empirical Study}
}