Search Results

Documents authored by Besenk, Mehmet


Artifact
Software
Lexcheck

Authors: Xingcheng Chen, Mehmet Besenk, and Andrea Stocco


Abstract

Cite as

Xingcheng Chen, Mehmet Besenk, Andrea Stocco. Lexcheck (Software, Source code). Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@misc{dagstuhl-artifact-27688,
   title = {{Lexcheck}}, 
   author = {Chen, Xingcheng and Besenk, Mehmet and Stocco, Andrea},
   note = {Software, version 1.0., swhId: \href{https://archive.softwareheritage.org/swh:1:dir:51739312cee13103280550edf2c8590475ed093d;origin=https://github.com/ast-fortiss-tum/lexcheck;visit=swh:1:snp:5175c5a7ec246f8653cb207867ddfba87ec01b71;anchor=swh:1:rev:2ca2bf4390e6c3de921d9f5e23d896144e4c9aeb}{\texttt{swh:1:dir:51739312cee13103280550edf2c8590475ed093d}} (visited on 2026-10-05)},
   url = {https://github.com/ast-fortiss-tum/lexcheck},
   doi = {10.4230/artifacts.27688},
}
Document
Technical Track Paper
Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study

Authors: Xingcheng Chen, Mehmet Besenk, and Andrea Stocco

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Task-specialized language models are increasingly integrated into software engineering workflows to support vertical-domain activities such as issue triaging, document classification, and automated analysis. Despite their adoption, there is limited empirical evidence on how to test their robustness and detect brittle behaviors under semantics-preserving input transformations. Aims. This paper investigates whether explainability-guided metamorphic testing can improve the effectiveness and validity of robustness testing for specialized language models compared to heuristic mutation strategies. Method. We conduct a large-scale empirical study of explanation-guided metamorphic testing across three datasets, four model architectures, and 20 testing configurations derived from combinations of attribution methods and mutation strategies. The evaluated configurations combine attribution-based token prioritization, LLM-driven mutation, and automated semantic verification to generate linguistically valid test variants. We assess failure discovery capability, semantic validity, and testing efficiency against heuristic baselines. Results. Explanation-guided metamorphic testing generates 2.30× more verified failure-inducing test cases than heuristic mutation strategies. Semantic verification substantially improves mutation validity and achieves high label-preservation precision among gate-accepted variants according to human annotation. The study further reveals systematic shortcut behaviors across models, including over-reliance on named entities and formatting cues. Conclusions. The results provide evidence that explanation-guided metamorphic testing is an effective and practical approach for empirically evaluating the robustness of task-specialized language models used in vertical AI applications.

Cite as

Xingcheng Chen, Mehmet Besenk, and Andrea Stocco. Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 5:1-5:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{chen_et_al:LIPIcs.ESEM.2026.5,
  author =	{Chen, Xingcheng and Besenk, Mehmet and Stocco, Andrea},
  title =	{{Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{5:1--5:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.5},
  URN =		{urn:nbn:de:0030-drops-279739},
  doi =		{10.4230/LIPIcs.ESEM.2026.5},
  annote =	{Keywords: Metamorphic testing, explainable AI, specialized language models, vertical AI, attribution-guided testing, automated test generation, empirical software engineering}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail