Search Results

Documents authored by Nam, Gyumin


Document
Technical Track Paper
Hidden Dependencies in LLM-Enabled Code Generation: An Empirical Study of Regressions and Escaped Failures

Authors: Gyumin Nam and Geunseok Yang

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Regression testing has traditionally focused on source-code changes and well-specified software dependencies such as compiler versions, library versions, and operating-system platforms. However, LLM-enabled code-generation systems also depend on non-code artifacts (the model, the prompt template, and the retrieval corpus) that can change independently of the application code. Updates to any of these artifacts can silently regress previously passing tasks, and the lightweight partial-test oracles commonly used in continuous evaluation may not detect such failures. Despite the practical relevance, paired regression rates, measurement validity, and escaped-failure rates under controlled non-code dependency changes have received little empirical attention. Aims. This work empirically analyzes how changes to models, prompts, and retrieval data affect the behavior of LLM-enabled code-generation systems, framing the problem as dependency-aware regression testing rather than a model or retrieval comparison. The questions addressed are failure-rate sensitivity, paired-regression frequency under the answer-bearing canonical retrieval condition, the measurement-validity decomposition of canonical regression signals (answer-bearing baseline rescue vs. non-canonical retrieval sensitivity), and how often generations pass a lightweight oracle yet fail a stronger oracle, reported over all generations and, conditionally, among stronger-oracle failures. Method. We designed a controlled empirical study that treats models, prompts, and retrieval data as software dependencies of an LLM-enabled code-generation pipeline. Starting from a fixed baseline, we varied one dependency dimension at a time across three open-weight 7B-class code LLMs, three prompt templates, and four retrieval-data conditions (one being a synthetic stress test), evaluated on 120 CodeRAG-Bench HumanEval tasks under three random seeds (2,880 generations across eight conditions). For each (task, condition, seed) triple, we recorded syntactic validity and per-assertion execution outcomes to derive failure rate, paired regression rate, escaped-failure rate, and across-seed flakiness, and compared each changed condition to the baseline using paired McNemar tests. Results. In the answer-bearing canonical benchmark condition, removing the retrieved document regressed 11.67% of previously passing tasks (D₁, p = 1.22 × 10⁻⁴, stable under task-level Bonferroni correction). Replacing it with a different-task mismatched document regressed 9.17% (D₂, p = 9.77 × 10⁻⁴, significant before correction but not after; reported as an auxiliary stressor). A copy-template audit showed that the M₀ baseline output was a normalized byte-level copy of the canonical retrieved document in 331 of 360 generations, and 12 of 14 D₁-regressed tasks exhibited this property. A title-based self-excluded BM25 proxy and an answer-stripped D₀ check showed that most of this canonical signal was due to answer-bearing baseline rescue (BM25 D₁: 3 regressions / 3 recoveries; D₂: 3 / 4); thus, harmful non-canonical retrieval degradation was not established in this sample. The intentionally permissive O₁ probe let 36/360 D₁ and 26/360 D₂ generations pass despite failing O₂; an order-dependent first-three-substantive variant reduced these counts to 8/360 and 3/360. Model and prompt perturbations are reported as framework coverage checks rather than causal claims. Conclusions. The strongest apparent regression signal arose from the answer-bearing canonical retrieval condition, but validity analyses showed that this signal primarily reflected a benchmark-induced copy-template dependency rather than an established deployment retrieval risk. These findings support dependency-aware paired regression testing and layered oracles, while cautioning that canonical answer-bearing retrieval regression rates should be interpreted as benchmark-specific stress signals rather than deployment-risk estimates.

Cite as

Gyumin Nam and Geunseok Yang. Hidden Dependencies in LLM-Enabled Code Generation: An Empirical Study of Regressions and Escaped Failures. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 51:1-51:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{nam_et_al:LIPIcs.ESEM.2026.51,
  author =	{Nam, Gyumin and Yang, Geunseok},
  title =	{{Hidden Dependencies in LLM-Enabled Code Generation: An Empirical Study of Regressions and Escaped Failures}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{51:1--51:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.51},
  URN =		{urn:nbn:de:0030-drops-280195},
  doi =		{10.4230/LIPIcs.ESEM.2026.51},
  annote =	{Keywords: LLM-enabled code generation, regression testing, retrieval-augmented generation, test oracles, empirical software engineering}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail