,
Matteo Esposito
,
Rick Kazman
,
Valentina Lenarduzzi
Creative Commons Attribution 4.0 International license
Context. Generative AI coding agents are increasingly used to automate software maintenance and issue resolution. However, current evaluations mainly focus on test-passing behavior and provide limited insights into whether agents modify the same software entities selected by developers. Aim. This study investigates the accuracy of issue localization performed by agents across different software granularities. Method. We analyzed 2,441 issue-fixing commits from 10 large-scale Java projects and evaluated three agents, each combining the OpenCode harness with a different open-weight LLM. We compared agent-modified entities against human-implemented fixes at the package, class, and method levels using Accuracy, Precision, Recall, F1-score, and MCC. Results. Agents partially identified the software entities requiring modification, but localization performance strongly depended on software granularity. Agents achieved the strongest results at the package level, while performance progressively degraded at the class and method levels. GLM-5 consistently achieved the strongest localization performance, although practical differences among agents remained limited. Our findings showed a limitation of agents implementing functionally plausible yet structurally different packages, class or methods leading to consistently negative MCC values worsening with finer granularity. Conclusion. Current agents can often identify the general architectural region affected by an issue, but still struggle to precisely localize fine-grained implementation points. These findings highlight both the potential and the limitations of agents for issue localization and motivate larger-scale investigations on localization behavior, architectural impact, and long-term maintainability implications.
@InProceedings{coppola_et_al:LIPIcs.ESEM.2026.59,
author = {Coppola, Antonino and Esposito, Matteo and Kazman, Rick and Lenarduzzi, Valentina},
title = {{On Coding Agent Issue Localization Accuracy - An Exploratory Study}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {59:1--59:13},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.59},
URN = {urn:nbn:de:0030-drops-280271},
doi = {10.4230/LIPIcs.ESEM.2026.59},
annote = {Keywords: AI, Coding Agents, Empirical Study}
}