,
Gang Lv,
Yifeng Gou,
Qizhao Wang,
Bo Zhang,
Ming Zhang,
Qing Xu,
Xiaosong Ding,
Yu Shen,
Yusang Xiong,
Xinyao Xiao,
Xinyu Liu
Creative Commons Attribution 4.0 International license
Industrial LLM compliance-auditing systems are maintained through both expert policy and executable support chains. After deployment, changing standards, user appeals, expert reviews, and report variants expose failures in rule interpretation, parsing, evidence binding, context assembly, output contracts, and regression checks. We propose CoEvolve, a feedback-driven policy-harness maintenance workflow in which an LLM maintenance agent proposes coordinated updates, an independent verifier checks target repair, regression, stability, and semantic risks, and experts confirm releasable changes. We report a five-month industrial experience in a Chinese compliance-auditing service, covering 33 production feedback instances, 128 audited systems, 384 reports, and 24,192 checklist items. The evaluation compares four frozen workflow states formed by different maintenance organizations; it is not a controlled comparison of methods on identical maintenance inputs. On a 979-item monthly holdout benchmark, the CoEvolve replay state reached 93.5% accuracy, 4.1 percentage points above prompt/SOP maintenance at 89.4%. Manual co-maintenance reached 93.6% (916/979), while CoEvolve reached 915/979; their Wilson intervals overlap and the five-system effective sample is small. Because the replay input included visible historical manual artifacts, this result shows that CoEvolve reconstructed a comparable release state from those artifacts, not that it substituted for manual co-maintenance. Across replay tasks, CoEvolve externalized 4.97 candidate attempts per feedback instance on average. The experience supports governed policy-harness maintenance without retraining and highlights verifier-gate predicates and blocked-candidate records as transferable release-governance mechanisms.
@InProceedings{peng_et_al:LIPIcs.ESEM.2026.80,
author = {Peng, Yiyang and Lv, Gang and Gou, Yifeng and Wang, Qizhao and Zhang, Bo and Zhang, Ming and Xu, Qing and Ding, Xiaosong and Shen, Yu and Xiong, Yusang and Xiao, Xinyao and Liu, Xinyu},
title = {{CoEvolve: Feedback-Driven Policy-Harness Maintenance for Industrial LLM Auditing Systems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {80:1--80:18},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.80},
URN = {urn:nbn:de:0030-drops-280480},
doi = {10.4230/LIPIcs.ESEM.2026.80},
annote = {Keywords: LLM auditing, policy-harness maintenance, production feedback, verifier gates, software engineering in practice}
}