Search Results

Documents authored by Wei, Dan


Document
Technical Track Paper
How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study

Authors: Shanggui Zhan, Xingqi Wang, Dan Wei, and Xin Xiang

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. LLM-as-Judge is increasingly adopted in automated program repair (APR) to assess patch correctness, yet its validity as a measurement instrument remains unaudited: potential benchmark-identifier leakage, evidence-presentation bias, and explanation unreliability have not been systematically examined. Aims. We investigate whether LLM Judge produces unbiased, stable, and well-grounded patch correctness judgments across three validity dimensions: metadata leakage, evidence presentation, and explanation reliability. Method. On a balanced, paired dataset of 326 patches from 163 Defects4J bugs, we evaluate GPT-4o and DeepSeek-V3 across a 12-setting prompt matrix covering benchmark-identifying metadata exposure and evidence presentation variants. We further conduct a manual analysis of 60 GPT-4o explanations using a five-category failure taxonomy and three independent annotators. Results. Under sanitized conditions, GPT-4o achieves an accuracy of 77.3% (MCC = 0.575, FPR = 17.2%), while DeepSeek-V3 achieves 80.4% accuracy (MCC = 0.625, FPR = 22.1%), with no statistically significant difference between the two models. Exposing combined benchmark-identifying metadata yields a suggestive accuracy increase of 4.3 percentage points for GPT-4o and 1.8 percentage points for DeepSeek-V3. Declaring that all tests have passed does not significantly bias either model; in contrast, providing a developer reference patch leads to substantial and statistically significant improvements for both models. Explanation failures are concentrated in misclassified cases, with the false-positive quadrant showing a 93.3% failure rate, primarily driven by hallucinated evidence and missed edge cases. Conclusions. LLM Judge is useful as a triage aid but is insufficient as a standalone correctness oracle. Reliable deployment requires prompt sanitization, explicit false positive rate reporting, clear distinction between reference-assisted and standalone assessment, and treating LLM explanations as investigative hypotheses rather than self-validating justifications.

Cite as

Shanggui Zhan, Xingqi Wang, Dan Wei, and Xin Xiang. How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 21:1-21:20, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{zhan_et_al:LIPIcs.ESEM.2026.21,
  author =	{Zhan, Shanggui and Wang, Xingqi and Wei, Dan and Xiang, Xin},
  title =	{{How Reliable Is LLM-as-Judge for Patch Correctness Assessment? An Empirical Study}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{21:1--21:20},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.21},
  URN =		{urn:nbn:de:0030-drops-279898},
  doi =		{10.4230/LIPIcs.ESEM.2026.21},
  annote =	{Keywords: automated program repair, patch correctness assessment, LLM-as-Judge, empirical study}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail