,
Mashiat Amin Farin
,
Yasin Sazid
,
Ahmedul Kabir
Creative Commons Attribution 4.0 International license
Background. Generating behavioral models like state transition diagrams from natural language requirements presents challenges in requirements analysis and design. Traditional NLP and ML methods struggle to maintain the correct execution flow and structural consistency. Recent advances in LLMs offer new opportunities, though their effectiveness in generating behavioral models, particularly concerning prompting, retrieval, and repair strategies, remains largely unexamined. Aims. This paper evaluates how LLM-based generation strategies (zero-shot, one-shot, few-shot, RAG) affect state transition diagram quality (correctness, completeness, understandability, terminological alignment) and whether iterative repair improves validity. Method. We built a dataset of 80 requirement–diagram pairs from software engineering textbooks. We generated PlantUML code for each requirement using four prompting strategies across four open-source LLMs. To isolate the effect of different knowledge sources, we conducted a RAG analysis study comparing four retrieval corpus configurations (full, examples-only, rules-only, theory-only). We also assessed an iterative repair workflow using automated evaluation. We evaluated syntactic validity via PlantUML validation and structural validity via rule-based analysis, supplementing with human evaluation of quality dimensions. We produced 1,169 PlantUML outputs, of which 120 validated outputs were manually evaluated. Results. Few-shot prompting improved syntactic validity from 28.7% to 93.5% on average over zero-shot. Iterative repair significantly improved structural validity, especially for DeepSeek (from 22.2% to 88.9%), with most repairs succeeding within two iterations. Human evaluation showed understandability was rated higher than correctness, indicating a quality-perception gap. Conclusions. We recommend few-shot prompting with iterative repair for AI-assisted behavioral modeling. Our findings reveal critical trade-offs: retrieval improves syntax but risks hallucination, model architecture matters more than strategy, and automated validation cannot substitute for human correctness assessment. This is the first empirical study systematically evaluating prompting, retrieval, and repair in LLM-based UML state transition diagram generation from natural language.
@InProceedings{faria_et_al:LIPIcs.ESEM.2026.19,
author = {Faria, Mussammat Maimuna and Farin, Mashiat Amin and Sazid, Yasin and Kabir, Ahmedul},
title = {{Towards Reliable AI-Assisted Behavioral Modeling: Evaluating Prompting, Retrieval, and Repair Strategies for UML State Transition Diagram Generation}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {19:1--19:21},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.19},
URN = {urn:nbn:de:0030-drops-279873},
doi = {10.4230/LIPIcs.ESEM.2026.19},
annote = {Keywords: State Transition Diagrams, Behavioral Software Modeling}
}
archived version
archived version