Search Results

Documents authored by Cánovas Izquierdo, Javier Luis


Document
Technical Track Paper
Assessing the Substitution of LLM-Based Decision Components

Authors: Sergio Cobos, Robert Clarisó, and Javier Luis Cánovas Izquierdo

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
Background. Large Language Models (LLMs) are increasingly used in software pipelines as decision components that validate, triage, or score system inputs and outputs. When one such component is replaced, for example, to reduce cost or latency or to adapt to provider-side model updates, it is necessary to make sure that the replacement does not significantly alter system behavior. Aims. This paper presents a methodology for assessing the operational substitutability of LLM-based decision components under controlled conditions. We combine two complementary criteria: instance-level behavioral agreement, measured with Cohen’s κ, and the absence of statistically significant asymmetries in correctness outcomes, assessed with McNemar’s test. Method. We instantiate the methodology on three representative binary decision tasks used in software pipelines: safety classification, truthfulness assessment, and stereotype detection. Using models from multiple providers, we study substitutability by asking (1) whether replacement within the same provider preserves decision behavior, (2) whether replacement across providers preserves decision behavior, and (3) how sensitive the admissible substitution space is to the agreement threshold and significance level used in the analysis. Results. Substitutability varies substantially across tasks and model pairs. Safety admits the broadest substitution space, while truthfulness is the most constrained task within providers and stereotype detection across providers. Sensitivity analyses show that operational parameter choices change the size of the admissible substitution space, but not the overall qualitative pattern. Conclusions. When assessing replacement candidates for LLM-based decision components, sufficient behavioral agreement and symmetry in correctness outcomes can complement aggregate benchmark performance.

Cite as

Sergio Cobos, Robert Clarisó, and Javier Luis Cánovas Izquierdo. Assessing the Substitution of LLM-Based Decision Components. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 12:1-12:21, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{cobos_et_al:LIPIcs.ESEM.2026.12,
  author =	{Cobos, Sergio and Claris\'{o}, Robert and C\'{a}novas Izquierdo, Javier Luis},
  title =	{{Assessing the Substitution of LLM-Based Decision Components}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{12:1--12:21},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.12},
  URN =		{urn:nbn:de:0030-drops-279803},
  doi =		{10.4230/LIPIcs.ESEM.2026.12},
  annote =	{Keywords: LLM-based decision components, model substitutability, LLM-as-a-judge, behavioral agreement, model replacement, software validation}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail