Search Results

Documents authored by Ribeiro, Eugénio


Document
Complete Volume
OASIcs, Volume 144, SLATE 2026, Complete Volume

Authors: Fernando Batista, Eugénio Ribeiro, Ricardo Ribeiro, and André L. Santos

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
OASIcs, Volume 144, SLATE 2026, Complete Volume

Cite as

15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 1-312, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@Proceedings{batista_et_al:OASIcs.SLATE.2026,
  title =	{{OASIcs, Volume 144, SLATE 2026, Complete Volume}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{1--312},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026},
  URN =		{urn:nbn:de:0030-drops-274870},
  doi =		{10.4230/OASIcs.SLATE.2026},
  annote =	{Keywords: OASIcs, Volume 144, SLATE 2026, Complete Volume}
}
Document
Front Matter
Front Matter, Table of Contents, Preface, Conference Organization

Authors: Fernando Batista, Eugénio Ribeiro, Ricardo Ribeiro, and André L. Santos

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Front Matter, Table of Contents, Preface, Conference Organization

Cite as

15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 0:i-0:xii, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{batista_et_al:OASIcs.SLATE.2026.0,
  author =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  title =	{{Front Matter, Table of Contents, Preface, Conference Organization}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{0:i--0:xii},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.0},
  URN =		{urn:nbn:de:0030-drops-274863},
  doi =		{10.4230/OASIcs.SLATE.2026.0},
  annote =	{Keywords: Front Matter, Table of Contents, Preface, Conference Organization}
}
Document
Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification

Authors: Eugénio Ribeiro and Jorge Baptista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
The automatic assessment of text readability and the classification of texts by levels is essential for language education and language industries that rely on effective communication. This study explores cross-language automatic readability level classification using the levels defined by the Common European Framework of Reference for Languages (CEFR). We investigate the potential of using data in one language to improve classification performance in different languages and, thus, optimize the utilization of the limited labeled resources available for each language. We rely on a pre-trained multilingual Transformer-based language model, by fine-tuning it on annotated data in one language or in a combination of languages, and then assessing its ability to generalize even to unseen languages. In an additional scenario, we further fine-tune the models on data in the target language, to assess whether the models trained on data in different languages can capture generic information regarding text readability and then be further specialized to capture the specific characteristics of the target language. Our experiments covering the English, Dutch, and German languages revealed that direct generalization to unseen languages is challenging. However, when paired with data in the target language, multilingual data can be leveraged to capture cross-language aspects of text readability, leading to more robust and better-performing models.

Cite as

Eugénio Ribeiro and Jorge Baptista. Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 3:1-3:13, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{ribeiro_et_al:OASIcs.SLATE.2026.3,
  author =	{Ribeiro, Eug\'{e}nio and Baptista, Jorge},
  title =	{{Cross-Language Text Readability Assessment: Leveraging Multilingual Models for Improved Performance in CEFR-Level Classification}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{3:1--3:13},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.3},
  URN =		{urn:nbn:de:0030-drops-267016},
  doi =		{10.4230/OASIcs.SLATE.2026.3},
  annote =	{Keywords: Readability, Text Complexity, CEFR, Multilinguality}
}
Document
BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy

Authors: João Francisco Botas, Ana Rita Peixoto, and Eugénio Ribeiro

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
The Bloom Taxonomy provides a useful framework for describing the cognitive demands of Multiple Choice Questions (MCQs), yet automatically assigning Bloom levels remains challenging due to category overlap and the limited evidence of higher-order thinking in MCQs. This study aims to evaluate transformer-based models, including BERT and ModernBERT variants, to enhance their performance on a four-level Bloom classification task across multiple scenarios, while also conducting model interpretability and error analysis. Across 40 settings tested, BERT models consistently outperform a keyword-based baseline, highlighting the contextual and semantic representations for this task. Performance is broadly similar between BERT and ModernBERT, with only minor differences across architectures, while the inclusion of answer options yields only small gains, suggesting that the question stem alone contains most of the discriminative signal. Besides, cross-linguistic experiments show comparable results between the original English data and Portuguese, although translated data exhibits a slight performance degradation, likely due to direct translation noise and subtle syntactic shifts. Errors are concentrated between adjacent Bloom levels, and models perform best on the Remembering level. In contrast, the Applying level remains the most difficult class to predict, largely because of class imbalance and conceptual overlap. The findings indicate that our approach constitutes a stable pipeline for Bloom classification, but its performance remains constrained by the intrinsic ambiguity of the taxonomy and the structural characteristics of the available data.

Cite as

João Francisco Botas, Ana Rita Peixoto, and Eugénio Ribeiro. BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 4:1-4:17, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{botas_et_al:OASIcs.SLATE.2026.4,
  author =	{Botas, Jo\~{a}o Francisco and Peixoto, Ana Rita and Ribeiro, Eug\'{e}nio},
  title =	{{BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{4:1--4:17},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.4},
  URN =		{urn:nbn:de:0030-drops-267025},
  doi =		{10.4230/OASIcs.SLATE.2026.4},
  annote =	{Keywords: Bloom Taxonomy, Multiple Choice Questions, Transformers, Multi-Language, Natural Language Processing, Educational Assessment}
}
Document
Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese

Authors: Raimundo Juracy Campos Ferro Junior, Eugénio Ribeiro, and Leonardo Sampaio Rocha

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
This study investigates bias in static word embeddings applied to Pajubá, a dialect spoken within the Brazilian LGBTQIAP+ community. Four models from the NILC repository (Word2Vec, GloVe, FastText, and Wang2Vec) were evaluated across two dimensions: representational capacity and polarity analysis. The vocabulary coverage test revealed that approximately 74% of dialectal terms are present in the models' vector spaces. The RND and SC-WEAT tests consistently point toward negative associations for identity terms across all models, with the SC-WEAT further suggesting that reappropriated dialect terms carry their standard Portuguese positive valence into the embedding space, while explicit identitarian markers remain negatively encoded. An adjectivation test confirms stereotypical associations, including the linkage of travesti with criminality and the pathologisation of gay. These findings suggest that static word embeddings fail to capture the sociolinguistic complexity of Pajubá and systematically reflect the dominant and often discriminatory discourses present in the training corpora.

Cite as

Raimundo Juracy Campos Ferro Junior, Eugénio Ribeiro, and Leonardo Sampaio Rocha. Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 7:1-7:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{ferrojunior_et_al:OASIcs.SLATE.2026.7,
  author =	{Ferro Junior, Raimundo Juracy Campos and Ribeiro, Eug\'{e}nio and Rocha, Leonardo Sampaio},
  title =	{{Polarizations in Static Word Embeddings: Investigating Bias in Marginalized Dialects of Portuguese}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{7:1--7:15},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.7},
  URN =		{urn:nbn:de:0030-drops-267055},
  doi =		{10.4230/OASIcs.SLATE.2026.7},
  annote =	{Keywords: word embeddings, bias detection, Portuguese, LGBTQIAP+, pajub\'{a}, NLP, lexical association, static embeddings}
}
Document
Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs

Authors: Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Addressing Online Hate Speech (OHS) in Portuguese faces significant challenges, including a lack of comprehensive, annotated data and classification frameworks, effective detection tools, and the high computational cost of Large Language Models (LLMs). This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs (ranging from 3B to 27B parameters) through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project. Findings indicate that models such as Mistral Small 3.2 24B, Gemma 3 27B, and Phi-4 achieve competitive performance, balancing precision and recall, whilst remaining computationally affordable. While excelling in identifying direct and indirect hate speech, out-group derogation, and specific target groups, the models faced challenges in detecting subtle rhetorical devices and emotions. Analysis of the Cohen’s Kappa and F1-scores revealed moderate agreement with human annotators, highlighting the potential of optimised models to democratise OHS detection in resource-constrained settings, alongside the need for further refinement to bridge the gap in human-level nuance and improve generalisation. This study aims to broaden knowledge regarding the detection of OHS in Portuguese, by demonstrating a viable, competitive, and low-cost approach that does not rely on large-scale infrastructure or expensive proprietary models.

Cite as

Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista. Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 9:1-9:16, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{cardoso_et_al:OASIcs.SLATE.2026.9,
  author =	{Cardoso, Mauro and Ribeiro, Eug\'{e}nio and Batista, Fernando},
  title =	{{Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{9:1--9:16},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.9},
  URN =		{urn:nbn:de:0030-drops-267077},
  doi =		{10.4230/OASIcs.SLATE.2026.9},
  annote =	{Keywords: Hate Speech Detection, Zero-shot Classification, Lightweight Language Models}
}
Document
Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling

Authors: Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista

Published in: OASIcs, Volume 135, 14th Symposium on Languages, Applications and Technologies (SLATE 2025)


Abstract
This study analyzes the social media discourse of leading figures from Portugal’s far right party CHEGA, examining 10,323 posts on X (formerly Twitter) published between late 2019 and mid‑2024. Using BERTopic, 59 latent topics clustered into two main discursive dynamics were found: (1) ideological and public, and (2) party, electoral and parliamentary related. Within the first dynamic, we conducted a focused sub-analysis of themes related with identity, immigration and security narratives - topics that display posting peaks around electoral cycles, suggesting the strategic use of emotionally charged, identitarian frames for political mobilization. The model exhibits strong topic coherence and lexical diversity, indicating its robustness in extracting thematic structures from politically polarized microtexts. Nevertheless, our findings are constrained by source, the absence of interaction metrics, and the unmet need to link online discourse to offline events. This study demonstrates how computational topic modeling can reveal strategic communication patterns in far-right political discourse and underscores the need for cross-platform and interaction-level research to assess broader societal impact.

Cite as

Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista. Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling. In 14th Symposium on Languages, Applications and Technologies (SLATE 2025). Open Access Series in Informatics (OASIcs), Volume 135, pp. 12:1-12:16, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2025)


Copy BibTex To Clipboard

@InProceedings{cardoso_et_al:OASIcs.SLATE.2025.12,
  author =	{Cardoso, Mauro and Ribeiro, Eug\'{e}nio and Batista, Fernando},
  title =	{{Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling}},
  booktitle =	{14th Symposium on Languages, Applications and Technologies (SLATE 2025)},
  pages =	{12:1--12:16},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-387-4},
  ISSN =	{2190-6807},
  year =	{2025},
  volume =	{135},
  editor =	{Baptista, Jorge and Barateiro, Jos\'{e}},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2025.12},
  URN =		{urn:nbn:de:0030-drops-236929},
  doi =		{10.4230/OASIcs.SLATE.2025.12},
  annote =	{Keywords: Political Discourse, Topic Modeling, Far-Right, CHEGA (Portugal), Social Media}
}
Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail