65 Search Results for "Batista, Fernando"


Volume

OASIcs, Volume 144

15th Symposium on Languages, Applications and Technologies (SLATE 2026)

SLATE 2026, Lisbon, Portugal, June 25-26, 2026

Editors: Fernando Batista, Eugénio Ribeiro, Ricardo Ribeiro, and André L. Santos

Volume

OASIcs, Volume 74

8th Symposium on Languages, Applications and Technologies (SLATE 2019)

SLATE 2019, June 27-28, 2019, Coimbra, Portugal

Editors: Ricardo Rodrigues, Jan Janoušek, Luís Ferreira, Luísa Coheur, Fernando Batista, and Hugo Gonçalo Oliveira

Document
Complete Volume
OASIcs, Volume 144, SLATE 2026, Complete Volume

Authors: Fernando Batista, Eugénio Ribeiro, Ricardo Ribeiro, and André L. Santos

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
OASIcs, Volume 144, SLATE 2026, Complete Volume

Cite as

15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 1-312, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@Proceedings{batista_et_al:OASIcs.SLATE.2026,
  title =	{{OASIcs, Volume 144, SLATE 2026, Complete Volume}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{1--312},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026},
  URN =		{urn:nbn:de:0030-drops-274870},
  doi =		{10.4230/OASIcs.SLATE.2026},
  annote =	{Keywords: OASIcs, Volume 144, SLATE 2026, Complete Volume}
}
Document
Front Matter
Front Matter, Table of Contents, Preface, Conference Organization

Authors: Fernando Batista, Eugénio Ribeiro, Ricardo Ribeiro, and André L. Santos

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Front Matter, Table of Contents, Preface, Conference Organization

Cite as

15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 0:i-0:xii, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{batista_et_al:OASIcs.SLATE.2026.0,
  author =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  title =	{{Front Matter, Table of Contents, Preface, Conference Organization}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{0:i--0:xii},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.0},
  URN =		{urn:nbn:de:0030-drops-274863},
  doi =		{10.4230/OASIcs.SLATE.2026.0},
  annote =	{Keywords: Front Matter, Table of Contents, Preface, Conference Organization}
}
Document
Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa

Authors: Pedro Moura, Inês Gama, Fernando Batista, and António Lopes

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Information retrieval for question-answering remains underexplored in specialised domains and under-resourced language variants such as European Portuguese. Existing benchmarks largely target general-domain English data and document-centric retrieval, failing to capture the semantic alignment required for linguistic consultation tasks over curated question–answer (QA) pairs. We address this gap by introducing a controlled evaluation framework for retrieval over the "Ciberdúvidas da Língua Portuguesa" corpus, comprising 29,145 expert-validated QA entries. Our approach systematically analyses the interaction between indexing strategies, encoder models, and retrieval paradigms, while modelling real-world query variability through a paraphrase-based benchmark of 600 queries across five user profiles, manually validated by a professional linguist to ensure semantic fidelity. Experiments show that dense retrieval with an IR-optimised monolingual encoder significantly outperforms both sparse (BM25) and hybrid methods, achieving a Mean Reciprocal Rank (MRR) of 0.93. Notably, hybrid retrieval underperforms due to lexical mismatch interference, challenging prevailing assumptions in the literature. Our contributions include a novel benchmark framework for linguistic QA retrieval, empirical evidence supporting monolingual IR-specialised models, and insights into retrieval robustness under paraphrastic variation, enabling improved QA systems for specialised and low-resource environments.

Cite as

Pedro Moura, Inês Gama, Fernando Batista, and António Lopes. Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberdúvidas Da Língua Portuguesa. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 1:1-1:17, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{moura_et_al:OASIcs.SLATE.2026.1,
  author =	{Moura, Pedro and Gama, In\^{e}s and Batista, Fernando and Lopes, Ant\'{o}nio},
  title =	{{Optimising Retrieval for Linguistic Question-Answering in European Portuguese: A Benchmark on Ciberd\'{u}vidas Da L{\'\i}ngua Portuguesa}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{1:1--1:17},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.1},
  URN =		{urn:nbn:de:0030-drops-266997},
  doi =		{10.4230/OASIcs.SLATE.2026.1},
  annote =	{Keywords: Information Retrieval, Question-Answering, European Portuguese, Sentence Encoders, Natural Language Processing}
}
Document
Evaluating the Prosodic Diversity of TTS Models for L2 Prosody Assessment

Authors: Mariana Julião, Alberto Abad, and Helena Moniz

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Recent advances in text-to-speech (TTS) have led to synthetic speech that is often indistinguishable from natural speech at the level of individual utterances. However, it remains unclear whether such systems reproduce the prosodic variability observed in natural speech in a large native-speaker population. This question is particularly relevant for applications in computer-assisted language learning (CALL), as variability is a central property of prosody and a prerequisite for robust assessment. In this work, we investigate whether TTS can approximate the distribution of prosodic patterns found in native speech, and whether it can serve as a reference for evaluating second language (L2) productions. To this end, we first compare different TTS models to native speakers, and then compare L2 speakers to synthetic speech by the model previously seen as the closest to native speakers. Our results show that regardless of TTS achieving high perceptual quality, its prosodic variability substantially differs from that of native speakers. As a consequence, comparisons between L2 speech and TTS-based references reveal both the potential and the limitations of using synthetic speech for prosody assessment.

Cite as

Mariana Julião, Alberto Abad, and Helena Moniz. Evaluating the Prosodic Diversity of TTS Models for L2 Prosody Assessment. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 2:1-2:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{juliao_et_al:OASIcs.SLATE.2026.2,
  author =	{Juli\~{a}o, Mariana and Abad, Alberto and Moniz, Helena},
  title =	{{Evaluating the Prosodic Diversity of TTS Models for L2 Prosody Assessment}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{2:1--2:13},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.2},
  URN =		{urn:nbn:de:0030-drops-267007},
  doi =		{10.4230/OASIcs.SLATE.2026.2},
  annote =	{Keywords: Speech synthesis, TTS, prosody, L1, L2}
}
Document
Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification

Authors: Eugénio Ribeiro and Jorge Baptista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
The automatic assessment of text readability and the classification of texts by levels is essential for language education and language industries that rely on effective communication. This study explores cross-language automatic readability level classification using the levels defined by the Common European Framework of Reference for Languages (CEFR). We investigate the potential of using data in one language to improve classification performance in different languages and, thus, optimize the utilization of the limited labeled resources available for each language. We rely on a pre-trained multilingual Transformer-based language model, by fine-tuning it on annotated data in one language or in a combination of languages, and then assessing its ability to generalize even to unseen languages. In an additional scenario, we further fine-tune the models on data in the target language, to assess whether the models trained on data in different languages can capture generic information regarding text readability and then be further specialized to capture the specific characteristics of the target language. Our experiments covering the English, Dutch, and German languages revealed that direct generalization to unseen languages is challenging. However, when paired with data in the target language, multilingual data can be leveraged to capture cross-language aspects of text readability, leading to more robust and better-performing models.

Cite as

Eugénio Ribeiro and Jorge Baptista. Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 3:1-3:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{ribeiro_et_al:OASIcs.SLATE.2026.3,
  author =	{Ribeiro, Eug\'{e}nio and Baptista, Jorge},
  title =	{{Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{3:1--3:13},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.3},
  URN =		{urn:nbn:de:0030-drops-267016},
  doi =		{10.4230/OASIcs.SLATE.2026.3},
  annote =	{Keywords: Readability, Text Complexity, CEFR, Multilinguality}
}
Document
BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy

Authors: João Francisco Botas, Ana Rita Peixoto, and Eugénio Ribeiro

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
The Bloom Taxonomy provides a useful framework for describing the cognitive demands of Multiple Choice Questions (MCQs), yet automatically assigning Bloom levels remains challenging due to category overlap and the limited evidence of higher-order thinking in MCQs. This study aims to evaluate transformer-based models, including BERT and ModernBERT variants, to enhance their performance on a four-level Bloom classification task across multiple scenarios, while also conducting model interpretability and error analysis. Across 40 settings tested, BERT models consistently outperform a keyword-based baseline, highlighting the contextual and semantic representations for this task. Performance is broadly similar between BERT and ModernBERT, with only minor differences across architectures, while the inclusion of answer options yields only small gains, suggesting that the question stem alone contains most of the discriminative signal. Besides, cross-linguistic experiments show comparable results between the original English data and Portuguese, although translated data exhibits a slight performance degradation, likely due to direct translation noise and subtle syntactic shifts. Errors are concentrated between adjacent Bloom levels, and models perform best on the Remembering level. In contrast, the Applying level remains the most difficult class to predict, largely because of class imbalance and conceptual overlap. The findings indicate that our approach constitutes a stable pipeline for Bloom classification, but its performance remains constrained by the intrinsic ambiguity of the taxonomy and the structural characteristics of the available data.

Cite as

João Francisco Botas, Ana Rita Peixoto, and Eugénio Ribeiro. BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 4:1-4:17, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{botas_et_al:OASIcs.SLATE.2026.4,
  author =	{Botas, Jo\~{a}o Francisco and Peixoto, Ana Rita and Ribeiro, Eug\'{e}nio},
  title =	{{BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{4:1--4:17},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.4},
  URN =		{urn:nbn:de:0030-drops-267025},
  doi =		{10.4230/OASIcs.SLATE.2026.4},
  annote =	{Keywords: Bloom Taxonomy, Multiple Choice Questions, Transformers, Multi-Language, Natural Language Processing, Educational Assessment}
}
Document
Inter-Sentential Relations in Enhanced Lexicalized Meaning Representation (E-LMR)

Authors: Jorge Baptista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Abstract Meaning Representation (AMR) is a widely adopted framework for graph-based sentence-level semantics. However, its abstraction from surface form and its focus on intra-sentential structure limit its ability to capture discourse-level phenomena. Uniform Meaning Representation (UMR) addresses this limitation by introducing cross-sentential mechanisms, at the cost of increased representational complexity. We argue that inter-sentential relations can be modeled without abandoning lexical anchoring. Within Lexicalized Meaning Representation, we propose a programmatic extension grounded in three domains: (i) nominal and pronominal anaphora, (ii) temporal anchoring, and (iii) inter-sentential cohesion devices. We call this Enhanced LMR (E-LMR). We contend that these relations are not abstract add-ons, but are lexically realized and should be represented accordingly. Rather than a full annotation scheme, we introduce design principles and representation strategies illustrated with European Portuguese data. A lexically grounded approach improves interpretability, annotation consistency, and cross-linguistic robustness. We thus position E-LMR as a viable alternative within a modular semantic architecture where discourse relations are explicitly encoded while remaining tightly coupled to lexical form.

Cite as

Jorge Baptista. Inter-Sentential Relations in Enhanced Lexicalized Meaning Representation (E-LMR). In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 5:1-5:18, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{baptista:OASIcs.SLATE.2026.5,
  author =	{Baptista, Jorge},
  title =	{{Inter-Sentential Relations in Enhanced Lexicalized Meaning Representation (E-LMR)}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{5:1--5:18},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.5},
  URN =		{urn:nbn:de:0030-drops-267030},
  doi =		{10.4230/OASIcs.SLATE.2026.5},
  annote =	{Keywords: Inter-sentential relations, Lexicalized Meaning Representation, Nominal and Pronominal Anaphora Resolution, Temporal Anaphora Resolution, Inter-Sentential Cohesion Devices, European Portuguese}
}
Document
A Topic–Word Graph for Topic Modeling Visualization

Authors: Ana Rita Peixoto

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Topic Modeling (TM) is a widely used method for discovering hidden themes in unlabeled text collections, but the resulting topics are often difficult to interpret. Most existing visualizations present topics separately and do not show the words that link them, offering only a partial view of how topics relate to one another. This work introduces a graph‑based representation in which topics and words are modeled as connected nodes. This structure makes shared words explicit and allows users to observe how individual terms contribute to multiple thematic areas within the corpus. By encoding topic–word relationships as edges, the visualization reveals the distribution of these connections and provides a more detailed perspective than traditional topic‑centric displays. The graph includes four node types, defined by their role and degree (number of connections). The developed tool, topicwordgraph, offers two complementary visualizations: a main graph representation that distinguishes node types by color, and two bar charts that summarize topic overlap and word memberships. Together, these views help users understand how topics intertwine through shared vocabulary. By making the shared words between topics explicit, this tool offers a practical and effective way to interpret how topics are interconnected.

Cite as

Ana Rita Peixoto. A Topic–Word Graph for Topic Modeling Visualization. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 6:1-6:11, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{peixoto:OASIcs.SLATE.2026.6,
  author =	{Peixoto, Ana Rita},
  title =	{{A Topic–Word Graph for Topic Modeling Visualization}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{6:1--6:11},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.6},
  URN =		{urn:nbn:de:0030-drops-267041},
  doi =		{10.4230/OASIcs.SLATE.2026.6},
  annote =	{Keywords: Topic Modeling, Data Visualization, Graph‑Based Representation, Topic Interpretability, Text Mining}
}
Document
Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese

Authors: Raimundo Juracy Campos Ferro Junior, Eugénio Ribeiro, and Leonardo Sampaio Rocha

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
This study investigates bias in static word embeddings applied to Pajubá, a dialect spoken within the Brazilian LGBTQIAP+ community. Four models from the NILC repository (Word2Vec, GloVe, FastText, and Wang2Vec) were evaluated across two dimensions: representational capacity and polarity analysis. The vocabulary coverage test revealed that approximately 74% of dialectal terms are present in the models' vector spaces. The RND and SC-WEAT tests consistently point toward negative associations for identity terms across all models, with the SC-WEAT further suggesting that reappropriated dialect terms carry their standard Portuguese positive valence into the embedding space, while explicit identitarian markers remain negatively encoded. An adjectivation test confirms stereotypical associations, including the linkage of travesti with criminality and the pathologisation of gay. These findings suggest that static word embeddings fail to capture the sociolinguistic complexity of Pajubá and systematically reflect the dominant and often discriminatory discourses present in the training corpora.

Cite as

Raimundo Juracy Campos Ferro Junior, Eugénio Ribeiro, and Leonardo Sampaio Rocha. Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 7:1-7:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{ferrojunior_et_al:OASIcs.SLATE.2026.7,
  author =	{Ferro Junior, Raimundo Juracy Campos and Ribeiro, Eug\'{e}nio and Rocha, Leonardo Sampaio},
  title =	{{Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{7:1--7:15},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.7},
  URN =		{urn:nbn:de:0030-drops-267055},
  doi =		{10.4230/OASIcs.SLATE.2026.7},
  annote =	{Keywords: word embeddings, bias detection, Portuguese, LGBTQIAP+, pajub\'{a}, NLP, lexical association, static embeddings}
}
Document
Gender Bias Evaluation in English-Portuguese Automated Translation Outputs

Authors: Xiaolan Xu and Sara Mendes

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Automated translation (AT) plays a pivotal role in breaking down language barriers and fostering cross-cultural communication. However, gender bias in AT remains a pressing concern, particularly for language pairs where the source language generally does not have grammatical gender and the target language does. In this context, the present research aims to assess gender bias in English-Portuguese AT outputs. Two machine translation (MT) systems and two LLMs are evaluated in this study, examining the correlation between translated gender and societal stereotypes, and how the different models handle nouns whose gender is unknown. The outputs of two commercial MT systems (Google Translate and DeepL Translator) and two general-purpose LLMs (ChatGPT-4o and DeepSeek-V3), translating from English into European Portuguese, were analyzed. We used a subset of the WinoMT challenge set to assess gender rendering accuracy across three conditions: gender-defined nouns, gender-undefined nouns, and anaphoric pronoun resolution. Our main goal is therefore to evaluate the accuracy of gender information transmission and to understand the extent to which gender bias is present in AT outputs. Our results indicate that all four models struggle to faithfully convey gender information from source to target, with male-centric outputs predominating across conditions and MT systems showing a more pronounced bias. This work contributes a preliminary cross-system comparison for an underexplored language pair and lays the groundwork for larger-scale evaluations of gender-inclusive AT.

Cite as

Xiaolan Xu and Sara Mendes. Gender Bias Evaluation in English-Portuguese Automated Translation Outputs. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 8:1-8:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{xu_et_al:OASIcs.SLATE.2026.8,
  author =	{Xu, Xiaolan and Mendes, Sara},
  title =	{{Gender Bias Evaluation in English-Portuguese Automated Translation Outputs}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{8:1--8:14},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.8},
  URN =		{urn:nbn:de:0030-drops-267065},
  doi =		{10.4230/OASIcs.SLATE.2026.8},
  annote =	{Keywords: Gender Bias, Large Language Models (LLMs), Neural Machine Translation (NMT), English-Portuguese Translation}
}
Document
Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs

Authors: Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Addressing Online Hate Speech (OHS) in Portuguese faces significant challenges, including a lack of comprehensive, annotated data and classification frameworks, effective detection tools, and the high computational cost of Large Language Models (LLMs). This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs (ranging from 3B to 27B parameters) through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project. Findings indicate that models such as Mistral Small 3.2 24B, Gemma 3 27B, and Phi-4 achieve competitive performance, balancing precision and recall, whilst remaining computationally affordable. While excelling in identifying direct and indirect hate speech, out-group derogation, and specific target groups, the models faced challenges in detecting subtle rhetorical devices and emotions. Analysis of the Cohen’s Kappa and F1-scores revealed moderate agreement with human annotators, highlighting the potential of optimised models to democratise OHS detection in resource-constrained settings, alongside the need for further refinement to bridge the gap in human-level nuance and improve generalisation. This study aims to broaden knowledge regarding the detection of OHS in Portuguese, by demonstrating a viable, competitive, and low-cost approach that does not rely on large-scale infrastructure or expensive proprietary models.

Cite as

Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista. Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 9:1-9:16, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{cardoso_et_al:OASIcs.SLATE.2026.9,
  author =	{Cardoso, Mauro and Ribeiro, Eug\'{e}nio and Batista, Fernando},
  title =	{{Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{9:1--9:16},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.9},
  URN =		{urn:nbn:de:0030-drops-267077},
  doi =		{10.4230/OASIcs.SLATE.2026.9},
  annote =	{Keywords: Hate Speech Detection, Zero-shot Classification, Lightweight Language Models}
}
Document
Assessing LLMs for Culturally Sensitive Communication in European Portuguese

Authors: Alice Vieira, Raquel Amaro, and Lyndon Nixon

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
This paper investigates the ability of LLMs to identify culturally sensitive communication in European Portuguese, in comparison with human performance. The study addresses cross-cultural communication issues arising within a single language, including miscommunication, bias, framing, and stereotyping. The task is formulated as a multi-component problem combining detection, classification, span identification, and reformulation of problematic content. A dataset of 60 sentences, balanced between problematic and neutral instances, was constructed from real-world corpora and annotated by expert linguists. Three open-weight multilingual models were evaluated alongside human annotators with diverse backgrounds. Results show that LLMs achieve high accuracy in detecting problematic content, outperforming non-expert participants and approaching expert performance. However, both models and humans exhibit low agreement in the classification of communication issues, reflecting the conceptual overlap between categories. While models generate consistent reformulations, they show limitations in span precision and contextual interpretation. The findings highlight a gap between detection and interpretation capabilities and support the use of LLMs as assistive tools in cross-cultural communication tasks, particularly in politically sensitive contexts.

Cite as

Alice Vieira, Raquel Amaro, and Lyndon Nixon. Assessing LLMs for Culturally Sensitive Communication in European Portuguese. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 10:1-10:14, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{vieira_et_al:OASIcs.SLATE.2026.10,
  author =	{Vieira, Alice and Amaro, Raquel and Nixon, Lyndon},
  title =	{{Assessing LLMs for Culturally Sensitive Communication in European Portuguese}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{10:1--10:14},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.10},
  URN =		{urn:nbn:de:0030-drops-267085},
  doi =		{10.4230/OASIcs.SLATE.2026.10},
  annote =	{Keywords: LLM, culturally sensitive communication, evaluation}
}
Document
Short Paper
Developing an AI-Based Application for Linguistic Analysis of Deaf Learners' L2 Writing (Short Paper)

Authors: Aryane Santos Nogueira and Ana Célia Ribeiro Bizigato Portes

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
The teaching of written language as a second language (L2) to deaf students remains a critical challenge in bilingual education. This is also the case for written Portuguese in Brazil, where difficulties persist particularly in the analysis of students' written production and its use in pedagogical decision-making. This paper presents the ongoing development of a Minimum Viable Product (MVP) aimed at supporting the qualified linguistic analysis of texts produced by deaf learners. Rather than functioning as a correction tool, the system is conceived as a pedagogically oriented analytical assistant, designed to identify patterns related to formal aspects of writing and their functional-communicative impact. The proposed model is structured around three analytical subdimensions (FORM_PROD, FORM_IMP, and FORM_FUN) and implemented through a prompt-based architecture using a Large Language Model (LLM). Preliminary results from a pilot study with five learner texts suggest the feasibility of distinguishing between formal instability and communicative effectiveness, generating interpretable and pedagogically relevant outputs. These findings point to the potential of AI to support language teaching in bilingual contexts involving deaf learners.

Cite as

Aryane Santos Nogueira and Ana Célia Ribeiro Bizigato Portes. Developing an AI-Based Application for Linguistic Analysis of Deaf Learners' L2 Writing (Short Paper). In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 11:1-11:10, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{nogueira_et_al:OASIcs.SLATE.2026.11,
  author =	{Nogueira, Aryane Santos and Portes, Ana C\'{e}lia Ribeiro Bizigato},
  title =	{{Developing an AI-Based Application for Linguistic Analysis of Deaf Learners' L2 Writing}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{11:1--11:10},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.11},
  URN =		{urn:nbn:de:0030-drops-267098},
  doi =		{10.4230/OASIcs.SLATE.2026.11},
  annote =	{Keywords: Deaf education, Second language writing, Linguistic analysis, Generative AI, Natural Language Processing}
}
  • Refine by Type
  • 63 Document/PDF
  • 2 Document/HTML
  • 2 Volume

  • Refine by Publication Year
  • 24 2026
  • 2 2025
  • 1 2024
  • 1 2023
  • 2 2022
  • Show More...

  • Refine by Author
  • 20 Batista, Fernando
  • 9 Ribeiro, Ricardo
  • 7 Ribeiro, Eugénio
  • 4 Coheur, Luísa
  • 4 Gonçalo Oliveira, Hugo
  • Show More...

  • Refine by Series/Journal
  • 2 LIPIcs
  • 61 OASIcs

  • Refine by Classification
  • 24 Computing methodologies → Natural language processing
  • 6 Computing methodologies → Language resources
  • 5 Software and its engineering → Domain specific languages
  • 4 Computing methodologies → Machine learning
  • 4 Human-centered computing → Natural language interfaces
  • Show More...

  • Refine by Keyword
  • 5 Natural Language Processing
  • 4 Portuguese Language
  • 3 Ontology
  • 3 Twitter
  • 2 Conference Organization
  • Show More...

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail