Deep Learning Approaches for Hate Speech Detection in Resource-Constrained Social Media Texts: A Comparative Evaluation

Authors

  • Yosia Immanuel Bastian Department of Informatics Engineering, Buddhi Dharma University, Indonesia
  • Aditiya Hermawan Department of Informatics Engineering, Buddhi Dharma University, Indonesia https://orcid.org/0000-0002-3372-8562
  • Ardiane Rossi Kurniawan Maranto Department of Information Systems, Buddhi Dharma University, Indonesia
  • Benny Daniawan Department of Information Systems, Buddhi Dharma University, Indonesia
  • Junaedi Junaedi Department of Information Systems, Buddhi Dharma University, Indonesia

DOI:

https://doi.org/10.47852/bonviewAIA62027379

Keywords:

BERT, deep learning, hate speech detection, low-resource language, BiLSTM

Abstract

Hate speech that occurred in social media platforms presents critical safety challenges, especially in language environments with minimal source where text is often written in an informal form, abbreviated, and context dependent. This paper further provides empirical evidence regarding the ability of the BERT (Bidirectional Encoder Representations from Transformers) model in detecting hate speech in a rough text environment and comparing directly with other deep learning techniques. Three Bidirectional Long Short–Term Memory (BiLSTM) variants using Word2Vec, FastText, and TF-IDF representations were compared with BERT, which employs contextual language representations. The data is separated by using a train-test ratio of 80:20, and the performance is evaluated according to their accuracy, precision, recall, F1-score, and area under the curve (AUC). The result shows that BERT outperforms all variations of BiLSTM by achieving an accuracy of 87.18%, an F1-score of 87.08%, and an AUC matrix of 0.9244. According to the efficiency matters, BiLSTM and Term Frequency–Inverse Document Frequency (TF-IDF) are the most ineffective for their uneven classification distribution, while BiLSTM with Word2Vec and FastText shows moderate effectivity. This finding conclusively demonstrates the benefit of transformer-based models to capture the subtleties of language with noisy textual data, which is commonly found in low-resource settings. To conclude the above discussion, this finding suggests that BERT has great potential as a multilingual content moderation tool to be applied in informal and unstructured digital environments.

 

Received: 25 August 2025 | Revised: 29 January 2026 | Accepted: 23 June 2026

 

Conflicts of Interest

The authors declare that they have no conflicts of interest to this work.

 

Data Availability Statement

The data that support the findings of this study are openly available in GitHub at https://github.com/okkyibrohim/id-multi-label-hate-speech-and-abusive-language-detection, reference number [29].

 

Author Contribution Statement

Yosia Immanuel Bastian: Conceptualization, Methodology, Formal analysis, Resources, Data curation. Aditiya Hermawan: Software, Supervision, Project administration, Writing – original draft, Writing – review & editing, Visualization. Ardiane Rossi Kurniawan Maranto: Validation, Investigation. Benny Daniawan: Validation, Investigation. Junaedi Junaedi: Validation, Investigation.


Downloads

Published

2026-07-20

Issue

Section

Research Article

How to Cite

Bastian, Y. I., Hermawan, A., Maranto, A. R. K., Daniawan, B., & Junaedi, J. (2026). Deep Learning Approaches for Hate Speech Detection in Resource-Constrained Social Media Texts: A Comparative Evaluation. Artificial Intelligence and Applications. https://doi.org/10.47852/bonviewAIA62027379