Cross-modal attention fusion using vision transformers for robust student attentiveness estimation

International Journal of Informatics and Communication Technology

Cross-modal attention fusion using vision transformers for robust student attentiveness estimation

Abstract

Automated student attentiveness estimation is a fundamental component of intelligent e-learning systems and adaptive classroom analytics. Traditional convolutional and recurrent architectures often struggle to model long-range temporal dependencies and complex inter-modal relationships inherent in engagement behavior. To address these limitations, this paper proposes a cross-modal attention fusion framework built upon a vision transformer (ViT) backbone for robust student attentiveness estimation. The proposed architecture leverages patch-based visual encoding through a ViT to capture global spatial dependencies, while behavioral cues such as gaze direction, head pose, and blink dynamics are embedded into a shared latent representation space. A cross-modal multi-head attention mechanism is introduced to dynamically learn interactions between visual and behavioral modalities, replacing static weighted fusion strategies. Temporal dynamics are modeled using a Transformer encoder, enabling effective long-range sequence modeling without recurrent dependencies. Experimental evaluation on a benchmark attentiveness dataset demonstrates superior performance compared to CNN–LSTM-based models, achieving improved accuracy, F1 score, and robustness under challenging lighting and occlusion conditions. Ablation studies validate the contribution of cross-modal attention and transformer-based temporal modeling. The proposed framework maintains real-time feasibility while significantly enhancing discriminative capability.

Discover Our Library

Embark on a journey through our expansive collection of articles and let curiosity lead your path to innovation.

Explore Now
Library 3D Ilustration