,
Boualem Benatallah
,
Silvana Togneri MacMahon
Creative Commons Attribution 4.0 International license
Background. Users can express the same request in many ways when interacting with LLM agents, particularly non-native English speakers whose writing reflects linguistic patterns transferred from their first language (L1). Existing benchmarks for evaluating API usage in LLM-based software systems typically rely on standardized English and overlook linguistic diversity. At the same time, collecting such data from users with diverse linguistic backgrounds is costly and difficult to scale. Aims. This study proposes L1-AUG, an L1-aware utterance generation method, and investigates three research questions: whether LLMs can generate API-calling utterances reproducing linguistic patterns of non-native English speakers; whether L1-aware utterance generation leads to higher linguistic diversity; and whether API-calling utterances that simulate non-native English linguistic patterns are more challenging for LLM agents to solve. Results. L1-AUG can generate utterances reflecting linguistic patterns extracted from essays written by non-native English speakers; L1-AUG generates benchmarks with higher linguistic diversity; and LLM agents achieve lower performance on the benchmark created using L1-AUG. Conclusions. Conditioning benchmark generation on linguistic patterns extracted from non-native English essays produces benchmarks that better reflect users with diverse linguistic backgrounds. Our results indicate that LLM agents struggle more with such inputs, suggesting that training on linguistically diverse datasets may be necessary to improve robustness in API-calling tasks.
@InProceedings{gaboardidossantos_et_al:LIPIcs.ESEM.2026.23,
author = {Gaboardi dos Santos, Vitor and Benatallah, Boualem and MacMahon, Silvana Togneri},
title = {{Beyond Standard English: L1-Aware Benchmarks for Evaluating API-Calling in LLM Agents}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {23:1--23:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.23},
URN = {urn:nbn:de:0030-drops-279916},
doi = {10.4230/LIPIcs.ESEM.2026.23},
annote = {Keywords: Software Testing Benchmarks, LLM-based Software Agents, Linguistic Diversity, L1 Transfer Patterns}
}
archived version