,
Ana Rita Peixoto
,
Eugénio Ribeiro
Creative Commons Attribution 4.0 International license
The Bloom Taxonomy provides a useful framework for describing the cognitive demands of Multiple Choice Questions (MCQs), yet automatically assigning Bloom levels remains challenging due to category overlap and the limited evidence of higher-order thinking in MCQs. This study aims to evaluate transformer-based models, including BERT and ModernBERT variants, to enhance their performance on a four-level Bloom classification task across multiple scenarios, while also conducting model interpretability and error analysis. Across 40 settings tested, BERT models consistently outperform a keyword-based baseline, highlighting the contextual and semantic representations for this task. Performance is broadly similar between BERT and ModernBERT, with only minor differences across architectures, while the inclusion of answer options yields only small gains, suggesting that the question stem alone contains most of the discriminative signal. Besides, cross-linguistic experiments show comparable results between the original English data and Portuguese, although translated data exhibits a slight performance degradation, likely due to direct translation noise and subtle syntactic shifts. Errors are concentrated between adjacent Bloom levels, and models perform best on the Remembering level. In contrast, the Applying level remains the most difficult class to predict, largely because of class imbalance and conceptual overlap. The findings indicate that our approach constitutes a stable pipeline for Bloom classification, but its performance remains constrained by the intrinsic ambiguity of the taxonomy and the structural characteristics of the available data.
@InProceedings{botas_et_al:OASIcs.SLATE.2026.4,
author = {Botas, Jo\~{a}o Francisco and Peixoto, Ana Rita and Ribeiro, Eug\'{e}nio},
title = {{BERT-Based Classification of Cross-Linguistic MCQs According to the Bloom Taxonomy}},
booktitle = {15th Symposium on Languages, Applications and Technologies (SLATE 2026)},
pages = {4:1--4:17},
series = {Open Access Series in Informatics (OASIcs)},
ISBN = {978-3-95977-440-6},
ISSN = {2190-6807},
year = {2026},
volume = {144},
editor = {Batista, Fernando and Ribeiro, Eug\'{e}nio and Ribeiro, Ricardo and Santos, Andr\'{e} L.},
publisher = {Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
address = {Dagstuhl, Germany},
URL = {https://drops.dagstuhl.de/entities/document/10.4230/OASIcs.SLATE.2026.4},
URN = {urn:nbn:de:0030-drops-267025},
doi = {10.4230/OASIcs.SLATE.2026.4},
annote = {Keywords: Bloom Taxonomy, Multiple Choice Questions, Transformers, Multi-Language, Natural Language Processing, Educational Assessment}
}