Deep Learning Approaches for Hate Speech Detection in Resource-Constrained Social Media Texts: A Comparative Evaluation
DOI:
https://doi.org/10.47852/bonviewAIA62027379Keywords:
BERT, deep learning, hate speech detection, low-resource language, BiLSTMAbstract
Hate speech that occurred in social media platforms presents critical safety challenges, especially in language environments with minimal source where text is often written in an informal form, abbreviated, and context dependent. This paper further provides empirical evidence regarding the ability of the BERT (Bidirectional Encoder Representations from Transformers) model in detecting hate speech in a rough text environment and comparing directly with other deep learning techniques. Three Bidirectional Long Short–Term Memory (BiLSTM) variants using Word2Vec, FastText, and TF-IDF representations were compared with BERT, which employs contextual language representations. The data is separated by using a train-test ratio of 80:20, and the performance is evaluated according to their accuracy, precision, recall, F1-score, and area under the curve (AUC). The result shows that BERT outperforms all variations of BiLSTM by achieving an accuracy of 87.18%, an F1-score of 87.08%, and an AUC matrix of 0.9244. According to the efficiency matters, BiLSTM and Term Frequency–Inverse Document Frequency (TF-IDF) are the most ineffective for their uneven classification distribution, while BiLSTM with Word2Vec and FastText shows moderate effectivity. This finding conclusively demonstrates the benefit of transformer-based models to capture the subtleties of language with noisy textual data, which is commonly found in low-resource settings. To conclude the above discussion, this finding suggests that BERT has great potential as a multilingual content moderation tool to be applied in informal and unstructured digital environments.
Received: 25 August 2025 | Revised: 29 January 2026 | Accepted: 23 June 2026
Conflicts of Interest
The authors declare that they have no conflicts of interest to this work.
Data Availability Statement
The data that support the findings of this study are openly available in GitHub at https://github.com/okkyibrohim/id-multi-label-hate-speech-and-abusive-language-detection, reference number [29].
Author Contribution Statement
Yosia Immanuel Bastian: Conceptualization, Methodology, Formal analysis, Resources, Data curation. Aditiya Hermawan: Software, Supervision, Project administration, Writing – original draft, Writing – review & editing, Visualization. Ardiane Rossi Kurniawan Maranto: Validation, Investigation. Benny Daniawan: Validation, Investigation. Junaedi Junaedi: Validation, Investigation.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
