NLP-driven hate speech detection on TikTok: a case study from UIN Sunan Ampel Surabaya

Telecommunication Computing Electronics and Control

NLP-driven hate speech detection on TikTok: a case study from UIN Sunan Ampel Surabaya

Abstract

This study examines hate speech detection in TikTok comments using natural language processing (NLP) techniques within the student community of UIN Sunan Ampel Surabaya. A dataset of 10,000 comments associated with the hashtag #PBAKUINSA2023 was analyzed using a lexicon-based sentiment analysis approach implemented through the TextBlob library, combined with Indonesian text preprocessing techniques, including tokenization, normalization, stopword removal, and stemming using the Sastrawi library. The results indicate that the proposed approach achieved an accuracy of 0.85, with precision of 0.88, recall of 0.83, and an F1-score of 0.854. Most comments were classified as neutral, while 31.8% were positive, and only a small proportion were negative. These findings suggest that discussions related to campus activities tend to be neutral or supportive. However, the findings also reveal that sentiment polarity does not always directly correspond to hate speech, as certain harmful expressions may appear neutral in lexicon-based analysis. This limitation highlights the need for more context-aware approaches. Overall, the proposed method provides an efficient solution for monitoring online discourse in academic environments.

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now
Library 3D Ilustration