,
João Pascoal Faria
Creative Commons Attribution 4.0 International license
Modern automotive infotainment systems sit at the center of the vehicle’s digital ecosystem, yet their validation still relies heavily on manual testing, a process that is time-consuming, expensive, and increasingly incompatible with the pace of agile release cycles and over-the-air software updates. Traditional scripted automation offers only a partial remedy, as it creates tight coupling between test logic and implementation details, producing brittle suites with high maintenance overhead. Existing LLM-driven testing frameworks predominantly target web or mobile applications, while employing single-agent or dual-agent architectures that overload one or two models with perception, planning, action selection, and validation simultaneously, making them prone to hallucination-style failures and unproductive exploration loops when faced with the complexity of automotive infotainment interfaces. In this paper, we present ARIA (Autonomous Real-time Infotainment Assessment), a multi-agent framework that leverages Large Language Models to autonomously execute end-to-end test scenarios on Android-based infotainment systems through visual interface interaction, orchestrating a closed-loop pipeline of four specialized agents per interaction step, complemented by a dedicated report-generation stage. From single-sentence natural-language scenario descriptions alone, each specifying a navigation path, an action to perform, and an expected outcome to verify, it autonomously executes the corresponding interactions on the infotainment system and produces structured reports, reproducible action scripts, and visual evidence for each step. ARIA was evaluated in an industrial setting on a physical test environment running the Android-based infotainment system of a car manufacturer, executing 30 scenarios spanning diverse system functionalities. Of the 30 scenarios, 28 (93.3%) completed the full multi-agent pipeline and produced a verdict, while 2 terminated prematurely with execution errors. Of the 28 completed scenarios, 20 (71.4%) matched the ground truth. ARIA detected all 5 known functional defects in the test setup, so no genuine fault was ever passed as working; the 8 false positives among completed scenarios are attributable to LLM navigation and image-interpretation limitations and to unsupported interaction gestures. These findings demonstrate that multi-agent LLM architectures can autonomously execute end-to-end infotainment test scenarios in an industrial setting, while also exposing the precision challenges that a low false-positive tolerance imposes. A single-agent baseline ablation on the same scenarios confirms the value of the multi-agent decomposition: on the first pass, before revisitation with a stronger model masks the difference, the single agent produces a substantially higher false positive rate (72.0% versus 52.6%) due to the conflation of navigational difficulty with system failure. We separately report first-pass and post-revisitation results and quantify token consumption, LLM-call counts, and monetary cost per scenario for both pipelines, and repeated execution of a representative subset of scenarios confirms that outcome stability correlates with scenario complexity, with fault detection remaining perfectly consistent across runs, indicating a path toward integrating visual test execution into continuous integration pipelines.
@InProceedings{azevedo_et_al:LIPIcs.ESEM.2026.83,
author = {Azevedo, Ant\'{o}nio and Lima, Bruno and Faria, Jo\~{a}o Pascoal},
title = {{ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems}},
booktitle = {20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
pages = {83:1--83:20},
series = {Leibniz International Proceedings in Informatics (LIPIcs)},
ISBN = {978-3-95977-450-5},
ISSN = {1868-8969},
year = {2026},
volume = {394},
editor = {Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.83},
URN = {urn:nbn:de:0030-drops-280517},
doi = {10.4230/LIPIcs.ESEM.2026.83},
annote = {Keywords: Testing, AI, LLM, Agentic, Infotainment, Autonomous, UI}
}