WVisionBERT-VL: a multimodal model architecture for toxicity classification on social media platforms using large language models

International Journal of Artificial Intelligence

WVisionBERT-VL: a multimodal model architecture for toxicity classification on social media platforms using large language models

Abstract

The increasing prevalence of toxic content on social media, conveyed through text, images, and videos, poses significant challenges for automated content moderation systems. Although prior studies have reported promising results in unimodal and bimodal settings, they often fail to capture implicit and contextual toxicity emerging from interactions across multiple modalities, particularly in non-English environments. This paper proposes WVisionBERT-VL, an end-to-end multimodal framework for toxicity detection that integrates text, image, and video modalities within a unified architecture. The proposed model incorporates modality-specific encoders, bidirectional multi-head cross attention (BMHCA) for cross-modal synchronization, and an adaptive fusion gate to dynamically balance modality contributions. A balanced multimodal dataset is constructed from social media platforms, including X, Instagram, and TikTok, and refined using a model-based labeling strategy with limited human-in-the-loop validation. Experimental results on a custom Indonesian dataset demonstrate strong in-domain performance, achieving an accuracy of 94.12%, a macro-F1 of 0.9407, and a receiver operating characteristic - area under the curve (ROC-AUC) of 0.9721, with robustness further validated through five-fold cross-validation. Cross-dataset evaluation highlights challenges related to domain shift, underscoring the need for future research on robust and domain-adaptive multimodal toxicity detection.

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now
Library 3D Ilustration