Search Results

Documents authored by Cardoso, Mauro


Document
Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs

Authors: Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista

Published in: OASIcs, Volume 144, 15th Symposium on Languages, Applications and Technologies (SLATE 2026)


Abstract
Addressing Online Hate Speech (OHS) in Portuguese faces significant challenges, including a lack of comprehensive, annotated data and classification frameworks, effective detection tools, and the high computational cost of Large Language Models (LLMs). This work evaluates the efficiency and competitiveness of seven smaller, more accessible LLMs (ranging from 3B to 27B parameters) through zero-shot classification with structured instructions, leveraging the expert-annotated test dataset and annotation scheme of the kNOwHATE project. Findings indicate that models such as Mistral Small 3.2 24B, Gemma 3 27B, and Phi-4 achieve competitive performance, balancing precision and recall, whilst remaining computationally affordable. While excelling in identifying direct and indirect hate speech, out-group derogation, and specific target groups, the models faced challenges in detecting subtle rhetorical devices and emotions. Analysis of the Cohen’s Kappa and F1-scores revealed moderate agreement with human annotators, highlighting the potential of optimised models to democratise OHS detection in resource-constrained settings, alongside the need for further refinement to bridge the gap in human-level nuance and improve generalisation. This study aims to broaden knowledge regarding the detection of OHS in Portuguese, by demonstrating a viable, competitive, and low-cost approach that does not rely on large-scale infrastructure or expensive proprietary models.

Cite as

Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista. Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs. In 15th Symposium on Languages, Applications and Technologies (SLATE 2026). Open Access Series in Informatics (OASIcs), Volume 144, pp. 9:1-9:16, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{cardoso_et_al:OASIcs.SLATE.2026.9,
  author =	{Cardoso, Mauro and Ribeiro, Eug\'{e}nio and Batista, Fernando},
  title =	{{Cost-Effective Hate Speech Detection in Portuguese Using Lightweight LLMs}},
  booktitle =	{15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
  pages =	{9:1--9:16},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-440-6},
  ISSN =	{2190-6807},
  year =	{2026},
  volume =	{144},
  editor =	{Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.9},
  URN =		{urn:nbn:de:0030-drops-267077},
  doi =		{10.4230/OASIcs.SLATE.2026.9},
  annote =	{Keywords: Hate Speech Detection, Zero-shot Classification, Lightweight Language Models}
}
Document
Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling

Authors: Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista

Published in: OASIcs, Volume 135, 14th Symposium on Languages, Applications and Technologies (SLATE 2025)


Abstract
This study analyzes the social media discourse of leading figures from Portugal’s far right party CHEGA, examining 10,323 posts on X (formerly Twitter) published between late 2019 and mid‑2024. Using BERTopic, 59 latent topics clustered into two main discursive dynamics were found: (1) ideological and public, and (2) party, electoral and parliamentary related. Within the first dynamic, we conducted a focused sub-analysis of themes related with identity, immigration and security narratives - topics that display posting peaks around electoral cycles, suggesting the strategic use of emotionally charged, identitarian frames for political mobilization. The model exhibits strong topic coherence and lexical diversity, indicating its robustness in extracting thematic structures from politically polarized microtexts. Nevertheless, our findings are constrained by source, the absence of interaction metrics, and the unmet need to link online discourse to offline events. This study demonstrates how computational topic modeling can reveal strategic communication patterns in far-right political discourse and underscores the need for cross-platform and interaction-level research to assess broader societal impact.

Cite as

Mauro Cardoso, Eugénio Ribeiro, and Fernando Batista. Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling. In 14th Symposium on Languages, Applications and Technologies (SLATE 2025). Open Access Series in Informatics (OASIcs), Volume 135, pp. 12:1-12:16, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2025)


Copy BibTex To Clipboard

@InProceedings{cardoso_et_al:OASIcs.SLATE.2025.12,
  author =	{Cardoso, Mauro and Ribeiro, Eug\'{e}nio and Batista, Fernando},
  title =	{{Portuguese Far-Right Discourse on Social Media: Insights from Topic Modeling}},
  booktitle =	{14th Symposium on Languages, Applications and Technologies (SLATE 2025)},
  pages =	{12:1--12:16},
  series =	{Open Access Series in Informatics (OASIcs)},
  ISBN =	{978-3-95977-387-4},
  ISSN =	{2190-6807},
  year =	{2025},
  volume =	{135},
  editor =	{Baptista, Jorge and Barateiro, Jos\'{e}},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2025.12},
  URN =		{urn:nbn:de:0030-drops-236929},
  doi =		{10.4230/OASIcs.SLATE.2025.12},
  annote =	{Keywords: Political Discourse, Topic Modeling, Far-Right, CHEGA (Portugal), Social Media}
}
Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail