Search Results

Documents authored by Tahsin, Noshin


Document
Emerging Results, Vision & Reflection Track Paper
Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering

Authors: Ryan Dang, Noshin Tahsin, and Thomas Zimmermann

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Threats to validity (TTV) are a core component of empirical research, yet software engineering studies fail to address them adequately. Published studies often omit validity discussions entirely, report generic concerns disconnected from the study’s actual methodology, or treat validity analysis as a procedural formality rather than a purposeful part of the research process. Aim. This study investigates whether large language models (LLMs) can generate threats to validity that are meaningful, study-specific, and novel, i.e., not reported by the original authors. Method. We conducted an empirical experiment using Gemini 2.5 Flash across 375 papers from ICSE 2025, withholding each paper’s TTV section prior to model input, ensuring the model reasoned from the study’s content rather than reproducing author-reported concerns. Generated threats were evaluated against a four-dimension rubric covering relevance, specificity, clarity, and mitigation quality, and compared against the original threats to determine novelty. All threats were then grouped into a taxonomy of seven categories to characterize the distribution of generated validity concerns. Results. The experiment produced 2,673 threats across the dataset, of which 97% were rated as relevant, specific, and clearly articulated, and 92% received high ratings for mitigation quality. Furthermore, 61% of the generated threats were absent from the original papers, suggesting that the model identified validity concerns beyond those reported by the original authors. Conclusions. LLMs can generate technically sound and previously unreported threats to validity at scale, demonstrating their capacity to support researchers in identifying overlooked methodological risks and strengthening the completeness of validity discussions, while augmenting rather than replacing researcher judgment in the process.

Cite as

Ryan Dang, Noshin Tahsin, and Thomas Zimmermann. Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 76:1-76:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{dang_et_al:LIPIcs.ESEM.2026.76,
  author =	{Dang, Ryan and Tahsin, Noshin and Zimmermann, Thomas},
  title =	{{Do LLMs Understand Validity? An Empirical Study of Machine‑Generated Threats to Validity in Software Engineering}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{76:1--76:14},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.76},
  URN =		{urn:nbn:de:0030-drops-280443},
  doi =		{10.4230/LIPIcs.ESEM.2026.76},
  annote =	{Keywords: Threats to Validity, Empirical Research, Large Language Models, LLMs, Software Engineering Research, Automated Analysis}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail