,
Bruno B. P. Cafeo
Creative Commons Attribution 4.0 International license
Background. LLM-based multi-agent systems (LLM-MAS) are increasingly applied to software engineering (SE) tasks. However, most systems are evaluated end-to-end, entangling architectural choices with model selection, prompting, and task-specific tooling. As a consequence, the effects of individual architectural factors remain empirically underexplored. Aims. We investigate how Communication Structure (Hierarchical vs. Centralized) and Coordination Strategy (Dynamic vs. Static) influence three outcome dimensions in LLM-MAS for Python refactoring: output generation and validity; operational cost and execution efficiency; and structural transformation. We also examine whether the two factors interact. Method. We conduct a controlled 2 × 2 factorial experiment under identical model, dataset, orchestration, prompting, and tool configurations, evaluated on 2,000 real-world Python refactoring instances drawn from open-source machine learning (ML) repositories. We measure output generation and validity through file generation, syntactic validity, and similarity-based acceptance; operational cost and execution efficiency through execution time, LLM calls, and token usage; and structural transformation through static code metrics. Results. Coordination Strategy is the dominant architectural factor: Static configurations achieve substantially higher output validity and lower operational cost across all token-based and call-based measures. Communication Structure exhibits a heterogeneous effect, strongly influencing call-based behavior but having limited impact on token consumption. The two factors interact significantly on all cost measures, with the strongest interaction observed on LLM calls. Among the evaluated configurations, Hierarchical+Static consistently provides the best tradeoff between validity and cost, whereas Centralized+Dynamic fails to produce acceptable output for a substantial fraction of instances. Conclusions. Coordination Strategy exerts more consistent effects than Communication Structure on both output validity and operational cost, while the interaction between the two factors is central to explaining architectural behavior. The study demonstrates how controlled factorial designs can isolate architectural effects in LLM-MAS research, and it provides empirical guidance for designing cost-efficient architectures with strong output validity for SE tasks.
@InProceedings{pumapucho_et_al:LIPIcs.ESEM.2026.37,
author = {Puma Pucho, Alexander and Cafeo, Bruno B. P.},
title = {{A Factorial Comparison of LLM-Based Multi-Agent Architectures for Automated Python Refactoring}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {37:1--37:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.37},
URN = {urn:nbn:de:0030-drops-280059},
doi = {10.4230/LIPIcs.ESEM.2026.37},
annote = {Keywords: Architectural Comparison, Code Refactoring, LLM-based Multi-Agent Systems, Software Maintenance Automation, Multi-Agent Architectures}
}