<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-05T21:32:51Z</responseDate>
  <request identifier="28048" metadataPrefix="oai_dc" verb="GetRecord">https://drops.dagstuhl.de/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:drops-oai.dagstuhl.de:28048</identifier>
        <datestamp>2026-10-05T06:44:06Z</datestamp>
        <setSpec>ddc:004</setSpec>
        <setSpec>open_access</setSpec>
      </header>
      <metadata>
        <oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
          <dc:title>CoEvolve: Feedback-Driven Policy-Harness Maintenance for Industrial LLM Auditing Systems</dc:title>
          <dc:creator>Peng, Yiyang</dc:creator>
          <dc:creator>Lv, Gang</dc:creator>
          <dc:creator>Gou, Yifeng</dc:creator>
          <dc:creator>Wang, Qizhao</dc:creator>
          <dc:creator>Zhang, Bo</dc:creator>
          <dc:creator>Zhang, Ming</dc:creator>
          <dc:creator>Xu, Qing</dc:creator>
          <dc:creator>Ding, Xiaosong</dc:creator>
          <dc:creator>Shen, Yu</dc:creator>
          <dc:creator>Xiong, Yusang</dc:creator>
          <dc:creator>Xiao, Xinyao</dc:creator>
          <dc:creator>Liu, Xinyu</dc:creator>
          <dc:subject>LLM auditing</dc:subject>
          <dc:subject>policy-harness maintenance</dc:subject>
          <dc:subject>production feedback</dc:subject>
          <dc:subject>verifier gates</dc:subject>
          <dc:subject>software engineering in practice</dc:subject>
          <dc:description>Industrial LLM compliance-auditing systems are maintained through both expert policy and executable support chains. After deployment, changing standards, user appeals, expert reviews, and report variants expose failures in rule interpretation, parsing, evidence binding, context assembly, output contracts, and regression checks. We propose CoEvolve, a feedback-driven policy-harness maintenance workflow in which an LLM maintenance agent proposes coordinated updates, an independent verifier checks target repair, regression, stability, and semantic risks, and experts confirm releasable changes. We report a five-month industrial experience in a Chinese compliance-auditing service, covering 33 production feedback instances, 128 audited systems, 384 reports, and 24,192 checklist items. The evaluation compares four frozen workflow states formed by different maintenance organizations; it is not a controlled comparison of methods on identical maintenance inputs. On a 979-item monthly holdout benchmark, the CoEvolve replay state reached 93.5% accuracy, 4.1 percentage points above prompt/SOP maintenance at 89.4%. Manual co-maintenance reached 93.6% (916/979), while CoEvolve reached 915/979; their Wilson intervals overlap and the five-system effective sample is small. Because the replay input included visible historical manual artifacts, this result shows that CoEvolve reconstructed a comparable release state from those artifacts, not that it substituted for manual co-maintenance. Across replay tasks, CoEvolve externalized 4.97 candidate attempts per feedback instance on average. The experience supports governed policy-harness maintenance without retraining and highlights verifier-gate predicates and blocked-candidate records as transferable release-governance mechanisms.</dc:description>
          <dc:publisher>Schloss Dagstuhl – Leibniz-Zentrum für Informatik</dc:publisher>
          <dc:contributor>Yiyang Peng and Gang Lv and Yifeng Gou and Qizhao Wang and Bo Zhang and Ming Zhang and Qing Xu and Xiaosong Ding and Yu Shen and Yusang Xiong and Xinyao Xiao and Xinyu Liu</dc:contributor>
          <dc:date>2026</dc:date>
          <dc:relation>Is Part Of LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)</dc:relation>
          <dc:type>InProceedings</dc:type>
          <dc:type>Text</dc:type>
          <dc:type>doc-type:ResearchArticle</dc:type>
          <dc:type>publishedVersion</dc:type>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>doi:10.4230/LIPIcs.ESEM.2026.80</dc:identifier>
          <dc:identifier>urn:nbn:de:0030-drops-280480</dc:identifier>
          <dc:identifier>https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.80</dc:identifier>
          <dc:language>eng</dc:language>
          <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
        </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
