,
Juri Di Rocco
,
Phuong T. Nguyen
,
Davide Di Ruscio
Creative Commons Attribution 4.0 International license
Background. Infrastructure-as-Code (IaC) scanners detect cloud misconfigurations in Terraform and other IaC languages before deployment, but repairing the flagged configurations remains largely manual. Recent Large Language Model (LLM)-based repair approaches can repair some findings, but may hallucinate unsupported constructs or suppress warnings without fixing the issue. Aims. We study whether tool grounding can improve LLM-based Terraform repair, and when a finding should be escalated because the required deployment-specific context is not available. Method. We present TerraRepair, a prototype of a tool-grounded LLM agent for Terraform repair with structured escalation. TerraRepair retrieves dependency context from Terraform references, consults the installed provider schema, and re-runs the scanner before returning a candidate repair. When the required context is absent, TerraRepair escalates instead of fabricating a plausible fix. Results. We evaluate our tool on two vulnerable-by-design Terraform repositories using two IaC security scanners, Checkov and Trivy, across AWS, Azure, and GCP. On the combined AWS benchmark, TerraRepair improves scanner-verified fix rates from 26.6% to 78.4% on Checkov and from 44.8% to 72.4% on Trivy, compared with a controlled one-shot baseline. It also reduces the baseline’s 44.8-73.6 percentage point (pp) claimed-vs-verified repair gap to under 5 pp. In a sampled semantic audit covering AWS only, 78.9% of TerraRepair’s scanner-verified AWS repairs are labeled as correct under a majority-vote protocol with two LLM judges and one author. Conclusions. These emerging results show that tool grounding can substantially improve scanner-verified LLM-based IaC repair on the studied benchmarks, while missing deployment-specific context remains the main knowledge boundary for full autonomy.
@InProceedings{mengistu_et_al:LIPIcs.ESEM.2026.68,
author = {Mengistu, Minase Mekete and Di Rocco, Juri and Nguyen, Phuong T. and Di Ruscio, Davide},
title = {{TerraRepair: A Tool-Grounded LLM Agent for Infrastructure-As-Code Repair}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {68:1--68:14},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.68},
URN = {urn:nbn:de:0030-drops-280363},
doi = {10.4230/LIPIcs.ESEM.2026.68},
annote = {Keywords: Infrastructure as Code, automated repair, large language models, cloud security}
}