,
Jorge Baptista
Creative Commons Attribution 4.0 International license
The automatic assessment of text readability and the classification of texts by levels is essential for language education and language industries that rely on effective communication. This study explores cross-language automatic readability level classification using the levels defined by the Common European Framework of Reference for Languages (CEFR). We investigate the potential of using data in one language to improve classification performance in different languages and, thus, optimize the utilization of the limited labeled resources available for each language. We rely on a pre-trained multilingual Transformer-based language model, by fine-tuning it on annotated data in one language or in a combination of languages, and then assessing its ability to generalize even to unseen languages. In an additional scenario, we further fine-tune the models on data in the target language, to assess whether the models trained on data in different languages can capture generic information regarding text readability and then be further specialized to capture the specific characteristics of the target language. Our experiments covering the English, Dutch, and German languages revealed that direct generalization to unseen languages is challenging. However, when paired with data in the target language, multilingual data can be leveraged to capture cross-language aspects of text readability, leading to more robust and better-performing models.
@InProceedings{ribeiro_et_al:OASIcs.SLATE.2026.3,
author = {Ribeiro, Eug\'{e}nio and Baptista, Jorge},
title = {{Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification}},
booktitle = {15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
pages = {3:1--3:13},
series = {Open Access Series in Informatics (OASIcs)},
ISBN = {978-3-95977-440-6},
ISSN = {2190-6807},
year = {2026},
volume = {144},
editor = {Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.3},
URN = {urn:nbn:de:0030-drops-267016},
doi = {10.4230/OASIcs.SLATE.2026.3},
annote = {Keywords: Readability, Text Complexity, CEFR, Multilinguality}
}